mirror of
https://github.com/fabro-sh/fabro.git
synced 2026-10-04 02:33:56 +00:00
452 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d1cc47324d
|
fix(agent): use raw sandbox reads for edits
Separate raw file reads from the line-numbered display API so apply_patch and edit_file operate on unformatted UTF-8 content. Keep read_file/read_many_files model-facing output numbered and cover regressions for prefix corruption. |
||
|
|
bf7ce485e1
|
feat(fabro-types): promote transcript primitives and extend agent event… (#357)
## Summary
This is the foundational step of the unified agent transcript
implementation: it establishes one canonical set of replay types in
`fabro-types` and threads them into the existing `agent.message`,
`agent.tool.started`, and `agent.tool.completed` event shapes — without
breaking any existing producers or consumers.
## What changed
**New `fabro-types::transcript` module** owns `ContentPart`,
`ImageData`, `AudioData`, `DocumentData`, `ThinkingData`, `ToolCall`,
`ToolResult`, `MessageKind`, `MessageSource`, `PairMessageRef`,
`TranscriptMessage`, and `MessageId`. These were previously defined in
`fabro-llm::types`; they now live at the canonical layer.
**`fabro-llm::types`** drops its local definitions and re-exports from
`fabro-types` so every existing `fabro_llm::types::*` import keeps
compiling without change.
**`AgentMessageProps`** gains an optional `message:
Option<TranscriptMessage>` field; `AgentToolStartedProps` gains
`tool_call`, `turn_id`, and `parent_message_id`;
`AgentToolCompletedProps` gains `tool_result` and `turn_id`. All new
fields use `#[serde(default, skip_serializing_if = "Option::is_none")]`
so existing stored events deserialize cleanly.
**All current event emitters** (`fabro-workflow/event/convert.rs`, demo
fixtures, test helpers) are updated to set the new fields to `None` —
this is a mechanical compatibility update; actual enrichment comes in
later tasks.
### Design decisions worth noting
- `MessageKind` captures LLM role semantics (system / user / reasoning /
agent); `MessageSource` captures audit provenance (steer, pair,
loop_detection, …). They are intentionally kept separate so a steering
message can be `kind=user, source=steer` without collapsing the
distinction.
- `TranscriptMessage` is named with the `Transcript` prefix specifically
to avoid import ambiguity with `fabro_agent::Message` and
`fabro_llm::types::Message`.
- `ProviderAnswer` and `ProviderReasoning` are included as
`MessageSource` variants so committed model outputs carry a first-class
audit label distinct from user-originated inputs.
- The new fields are additive-only; no narrow legacy fields were
removed. Consumer migration is a separate step.
### Fabro Details
<details>
<summary>Ran 9 stages in 47m 27s for $12.79</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 2m 6s | – | 0 |
| preflight_lint | 2m 21s | – | 0 |
| implement | 20m 6s | $8.53 | 0 |
| simplify_opus | 13m 36s | $2.68 | 0 |
| simplify_gpt | 4m 17s | $1.58 | 0 |
| verify | 4m 7s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **47m 27s** | **$12.79** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
Co-authored-by: Bryan Helmkamp <bryan@brynary.com>
|
||
|
|
a2e2cbc7ed
|
Add agent context observability events (memory, skills, MCP tools) (#356)
## Summary
Adds three new durable run events — `agent.memory.loaded`,
`agent.skills.discovered`, and `agent.skill.activated` — and enriches
`agent.mcp.ready` with names-only tool summaries. Consumers can now
reconstruct what memory, skills, and MCP tools were active for any agent
run by reading the event stream, without needing to inspect session
state.
### Plan Summary
- **`fabro-types`**: New prop structs (`AgentMemoryLoadedProps`,
`AgentSkillsDiscoveredProps`, `AgentSkillActivatedProps`,
`AgentMcpToolSummary`) and three new `EventBody` variants with canonical
dot-name serialization. `AgentMcpReadyProps.tools` uses
`#[serde(default, skip_serializing_if = "Vec::is_empty")]` for backwards
compatibility.
- **`fabro-agent/memory.rs`**: `discover_memory` now returns
`Vec<MemoryDocument>` carrying path, byte counts, and truncation flag
alongside content. The content itself is never put in any event payload.
- **`fabro-agent/types.rs`**: Adds `MemoryLoaded`, `SkillsDiscovered`,
`SkillActivated`, and enriched `McpServerReady` internal variants.
Removes `SkillExpanded` (replaced by `SkillActivated { source: Slash
}`). New variants are **not** classified as streaming noise, so they
persist.
- **`fabro-agent/session.rs`**: Emits `MemoryLoaded` before skills init,
`SkillsDiscovered` after skill discovery, and enriches `McpServerReady`
with summaries from `McpConnectionManager::tool_summaries_for_server`.
Slash expansion now emits `SkillActivated { source: Slash }` instead of
`SkillExpanded`.
- **`fabro-agent/skills.rs`**: `make_use_skill_tool` emits
`SkillActivated { source: Tool }` on successful lookup only.
- **`fabro-mcp/connection_manager.rs`**: New `tool_summaries_for_server`
returns sorted `(qualified_name, original_name)` pairs without leaking
descriptions or schemas.
- **`fabro-workflow/event/convert.rs` + `names.rs`**: Converts all new
agent events to their typed `fabro-types` props, including `visit`
injection. Removes dead `SkillExpanded` arm.
- **`docs/internal/events.md`**: Documents all new event shapes with
full property tables; notes that `agent.skill.expanded` is replaced.
### Key design decisions
- Both `MemoryLoaded` and `SkillsDiscovered` are emitted even when the
result is empty. This lets consumers distinguish "no memory/skills
found" from "event not yet reported."
- Memory file **contents are never included** in any event payload —
only `path`, `byte_count`, `loaded_bytes`, and `truncated`.
- `agent.mcp.ready` `tools` field is omitted from JSON when empty
(`skip_serializing_if`), preserving wire compatibility with existing
stored events.
- `SkillActivated` is persisted (not filtered as streaming noise),
unlike the former internal-only `SkillExpanded`.
### Fabro Details
<details>
<summary>Ran 9 stages in 57m 15s for $27.15</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 2m 17s | – | 0 |
| preflight_lint | 2m 30s | – | 0 |
| implement | 21m 18s | $16.27 | 0 |
| simplify_opus | 15m 21s | $6.52 | 0 |
| simplify_gpt | 10m 33s | $4.35 | 0 |
| verify | 4m 24s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **57m 15s** | **$27.15** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
|
||
|
|
0754f1ca4a
|
Add fabro_run_get read-only run inspection tool (#358)
## Summary
Adds a new `fabro_run_get` MCP tool that returns a run's summary,
resolved ID, projection, and pending questions without any mutation
capability. This separates read-only inspection from operational
control, allowing Ask Fabro sessions to inspect runs safely without
access to write operations.
## What changed and why
**New tool (`fabro-tool/src/get.rs`):** `FabroRunGetParams` /
`ValidatedRunGet` / `RunGetResult` follow the same
validation-and-dispatch pattern as other run tools. The implementation
resolves a selector, then fans out to three read-only API calls
(retrieve run, get state, list questions) and assembles them into a
single structured result.
**Tool registry and dispatch:** `FABRO_RUN_GET_TOOL_NAME` is exported
from `common.rs` and `lib.rs`, added to `TOOL_DEFINITIONS`, wired into
the MCP stdio server (`fabro-mcp-server/src/server.rs`), and dispatched
in the LLM agent executor (`fabro-workflow/src/handler/llm/api.rs`).
**Ask Fabro access policy (`fabro-server/.../sessions.rs`):** The
session now registers and allows only `fabro_run_events` +
`fabro_run_get` via the new `ASK_FABRO_RUN_TOOL_NAMES` constant.
`fabro_run_interact` is explicitly moved to the denied set, closing off
mutation from that session type. The policy match arm is refactored from
a hardcoded `|`-chain to a slice `contains` check so the constant is the
single source of truth.
**Docs:** `mcp.mdx` now lists `fabro_run_get` as the inspection tool and
redescribes `fabro_run_interact` as control-oriented.
**Backward compatibility:** `fabro_run_interact` (including `get` and
`get_questions` actions) is unchanged and still fully operational for
contexts that allow it.
### Plan Summary
- New `get.rs` module in `fabro-tool` with validation, async fetch, and
unit tests
- Constants + schema registration in `common.rs` / `lib.rs`
- MCP server and LLM dispatch branches added for
`FABRO_RUN_GET_TOOL_NAME`
- Ask Fabro session swaps `FABRO_RUN_INTERACT_TOOL_NAME` →
`FABRO_RUN_GET_TOOL_NAME` in registry and policy
- MCP integration tests: tool count constant, schema assertions, two new
end-to-end tests
(`mcp_get_resolves_selector_and_returns_summary_projection_and_questions`,
`mcp_get_rejects_blank_run_id_before_auth_or_network`)
- Public MCP docs updated
### Fabro Details
<details>
<summary>Ran 9 stages in 43m 30s for $13.42</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 2m 2s | – | 0 |
| preflight_lint | 2m 14s | – | 0 |
| implement | 19m 43s | $8.85 | 0 |
| simplify_opus | 5m 52s | $1.45 | 0 |
| simplify_gpt | 10m 16s | $3.12 | 0 |
| verify | 2m 53s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **43m 30s** | **$13.42** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
|
||
|
|
94c657b92e
|
feat: add run.checkpoint.skip_git_hooks to bypass Git commit hooks (#355)
## Summary
Adds an opt-in `skip_git_hooks` boolean to `[run.checkpoint]` that
causes Fabro-managed run-branch checkpoint commits to pass `--no-verify`
to `git commit`, bypassing local hooks such as `pre-commit` and
`commit-msg`. Defaults to `false`. Metadata-branch snapshots and Fabro
`[[run.hooks]]` are unaffected.
```toml
[run.checkpoint]
skip_git_hooks = true
```
### Plan Summary
- `RunCheckpointSettings` (dense, in `fabro-types`) gains
`skip_git_hooks: bool` with `#[serde(default)]`.
- `RunCheckpointLayer` (sparse, in `fabro-config`) gains
`skip_git_hooks: Option<bool>` so layered config can distinguish unset
from explicit `false`.
- `RunCheckpointLayer::combine` is refactored from a wholesale-replace
to field-level merging: `exclude_globs` keeps its existing replace-wins
semantics; `skip_git_hooks` uses `.or()` (highest-priority layer that
sets it wins).
- `resolve_checkpoint` resolves `None → false`.
- `git_checkpoint` / `checked_git_checkpoint` in `sandbox_git.rs` accept
a new `skip_git_hooks: bool` and append `--no-verify` when true.
- `parallel_branch_commit_cmd` (new helper in `handler/parallel.rs`)
replaces the inline format string and accepts the same flag.
- `GitState` carries `checkpoint_skip_git_hooks`;
`RunOptions::checkpoint_skip_git_hooks()` exposes it; `execute.rs` and
`git.rs` thread it through.
- OpenAPI schema, TypeScript API client, and docs are updated.
### Key design decisions
**Field-level merging in `combine`**: the previous
`RunCheckpointLayer::combine` replaced the whole struct when
`self.exclude_globs` was non-empty. The refactor keeps that same replace
rule for `exclude_globs` while adding independent `Option::or` merging
for `skip_git_hooks`, so the two fields don't interfere.
**`--no-verify` only on run-branch commits**: the flag is injected only
in the two Git commit paths Fabro controls for run-branch checkpoints.
Metadata-branch snapshots use `git2` and never fire local hooks
regardless of this setting.
### Fabro Details
<details>
<summary>Ran 9 stages in 48m 3s for $13.91</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 3s | – | 0 |
| preflight_compile | 2m 4s | – | 0 |
| preflight_lint | 2m 15s | – | 0 |
| implement | 24m 13s | $10.38 | 0 |
| simplify_opus | 10m 31s | $1.59 | 0 |
| simplify_gpt | 5m 8s | $1.93 | 0 |
| verify | 3m 4s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **48m 3s** | **$13.91** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
|
||
|
|
a42634cc54
|
Add event-sourced todo tools for OpenAI and Anthropic profiles (#353)
## Summary
Adds a shared todo/task engine behind two model-native tool surfaces,
with all mutations persisted as individual run events and replayed into
`RunProjection`.
### Plan Summary
- **New domain types** in `fabro-types`: `TodoStatus`, `TodoListKind`,
`TodoProjection`, `TodoListProjection`, and `todos_by_list` on
`RunProjection`.
- **New run events**: `todo.created`, `todo.updated`, `todo.deleted` —
mapped through Fabro's typed event pipeline and replayed by
`RunProjectionReducer`.
- **`TodoRuntime`** (`fabro-agent`): thread-safe in-memory projection
shared across tool closures within a profile instance; each mutation
emits the corresponding agent event.
- **`update_plan`** registered only in `OpenAiProfile`: reconciles
incoming steps by exact `step` text (sha256-derived ID), emitting
create/update/delete events to match the submitted plan.
- **`TaskCreate` / `TaskUpdate` / `TaskList`** registered only in
`AnthropicProfile`: numeric task IDs per list, metadata merge with
`null`-key deletion, `status: "deleted"` routes to `todo.deleted`.
- **Session identity threading**: `ToolContext` gains `session_id`,
`root_session_id`, `tool_call_id`, and `agent_event_emitter`;
`execute_tool_calls` threads these through to `execute_one_tool`;
`Session` tracks `root_session_id` and `spawn_agent` inherits it for
subagents.
- **Scoping**: OpenAI todos scope to `openai_plan:<session_id>`
(per-session); Anthropic todos scope to
`anthropic_tasks:<root_session_id>` (shared across subagents).
- **Web invalidation**: `todo.*` events invalidate `getRunState` and the
run events list; tested in `run-events.test.tsx`.
```mermaid
graph TB
subgraph OpenAI
UP[update_plan] -->|diff by step text| TR[TodoRuntime]
end
subgraph Anthropic
TC[TaskCreate] --> TR
TU[TaskUpdate] --> TR
TL[TaskList] -->|read-only snapshot| TR
end
TR -->|emit todo.created/updated/deleted| SE[SessionBoundEmitter]
SE --> EV[AgentEvent stream]
EV --> RP[RunProjection\ntodos_by_list]
```
### Key design decisions
- **Step identity by text, not position** (`update_plan`): a
sha256-derived ID from `list_id + step` means reordering without
renaming emits an update rather than a delete+create. Duplicate step
strings are rejected with a model-visible error because text is the
identity.
- **No plan-replace event**: the engine emits only individual mutation
events; bulk replacement is expressed as a set of create/update/delete
events produced by diffing the incoming plan against the projection
snapshot.
- **`TodoRuntime` per profile instance**: tools inside a single profile
share one runtime. OpenAI subagents each have their own `session_id` so
their plans are isolated; Anthropic subagents inherit `root_session_id`
so tasks are shared — matching upstream Codex/Claude behavior.
- **`AgentEventEmitter` trait on `ToolContext`**: a narrow interface
that lets tools publish typed events without taking a dependency on the
full `Emitter`. `SessionBoundEmitter` wraps `Emitter` and stamps
`session_id` + `tool_call_id` on each event.
### Fabro Details
<details>
<summary>Ran 9 stages in 93m 23s for $56.46</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 2m 6s | – | 0 |
| preflight_lint | 2m 20s | – | 0 |
| implement | 41m 30s | $38.21 | 0 |
| simplify_opus | 29m 2s | $11.78 | 0 |
| simplify_gpt | 14m 24s | $6.46 | 0 |
| verify | 3m 2s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **93m 23s** | **$56.46** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
Co-authored-by: Bryan Helmkamp <bryan@brynary.com>
|
||
|
|
95b45b5960
|
feat: wire Ask Fabro sidebar to real session API with run-control tools (#349)
## Summary
Ships the Ask Fabro sidebar on run pages end-to-end: the agent now has
live `fabro_run_interact` and `fabro_run_events` tools scoped to its
owning run, and the web sidebar talks to real session APIs instead of a
scripted adapter. The `?ask=1` prototype gate is dropped in favour of
server-reported `run.ask_fabro.available`.
## What changed and why
### Rust — run-control tools in Ask Fabro sessions (`fabro-server`,
`fabro-workflow`, `fabro-tool`)
**Tool registration** (`fabro-workflow`): `register_fabro_run_tools` is
now `pub`; a new `register_named_fabro_run_tools` variant accepts a name
allowlist so callers can register a subset without forking the catalog
loop. Unknown names are silently ignored.
**Run-scoped backend** (`fabro-tool`): `ClientBackend` gains a
`run_scope: Option<RunId>` field set via `.with_run_scope(run_id)`.
Every method checks the scope before delegating to the HTTP client,
returning an error before a network call is made. `list_store_runs`
returns a single-element vec of the owning run when scoped;
`resolve_run` rejects non-parseable selectors rather than forwarding
them.
**Session wiring** (`fabro-server`): `build_profile` now returns
`Box<dyn AgentProfile>` (mutably accessible) instead of `Arc`;
`build_agent_session` mints a same-run worker token, builds a
`ClientBackend::with_run_scope`, constructs `FabroRunToolServices`, and
calls `register_named_fabro_run_tools` for the two tools before freezing
into an `Arc`. `AppState::self_server_target()` reads the bound address
from the runtime daemon record for the loopback HTTP call.
**Approval gate**: `build_ask_fabro_tool_approval` now fast-paths
`fabro_run_interact` and `fabro_run_events` to `Ok(())`; all other tools
remain subject to the `ReadOnly` auto-approve check. File/shell tools
are still denied.
### Web — real session adapter and sidebar wiring (`fabro-web`)
**`ask-fabro-runtime.ts`** (new): a `ChatModelAdapter` that creates a
session lazily on the first turn (`sessionsApi.createRunSession`),
caches the session id in `sessionStorage` keyed by run id, and streams
turns via `streamSessionTurn`. `applyTurnEvent` maps `run.session.*` SSE
events to assistant-ui `ThreadAssistantMessagePart[]` incrementally
(text deltas, tool-call started/completed pairs). A 404 on stream clears
the cached id so the next turn starts fresh.
**`ask-fabro-sidebar.tsx`**: drops `scriptIndexRef`, `EMPTY_CHAT`, and
the scripted adapter import; accepts `runId` and `defaultModel` props;
constructs the real adapter via `createAskFabroAdapter`.
**`run-detail.tsx`**: removes `?ask=1` / `askEnabled`; reads
`run.ask_fabro.{available, default_model}` from the summary; always
renders an `AskFabroTriggerButton` (disabled with a tooltip when
unavailable); passes `runId` and `defaultModel` to `<AskFabroSidebar>`.
### Architecture
```mermaid
graph TB
Browser -->|SSE turn stream| SessionsHandler
SessionsHandler -->|spawn| AskFabroAgent
AskFabroAgent -->|fabro_run_interact\nfabro_run_events| ClientBackend
ClientBackend -->|HTTP + same-run\nworker token| RunsAPI[Runs API\n/runs/:id]
ClientBackend -->|run_scope check| ClientBackend
RunsAPI -->|403 cross-run| ClientBackend
```
### Design decisions
- **Same-run scoping is double-enforced**: the `ClientBackend` scope
check fires before the HTTP call; the worker token's run scope causes a
403 at the API layer if the check were somehow bypassed.
- **`build_profile` → `Box` not `Arc`**: the profile needs mutable
access for tool registration after construction, so the `Arc` wrapping
is deferred until registration is complete.
- **`sessionStorage` per-run**: one session is reused across sidebar
open/close cycles for the same run tab; a page reload or different run
always starts clean.
- **Mutating actions included**: `interact` exposes
start/cancel/steer/archive/answer. This is intentional per the locked
decisions; the worker-token scope prevents cross-run blast radius.
### Fabro Details
<details>
<summary>Ran 9 stages in 74m 52s for $44.11</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 2m 11s | – | 0 |
| preflight_lint | 2m 22s | – | 0 |
| implement | 38m 37s | $32.02 | 0 |
| simplify_opus | 17m 40s | $6.55 | 0 |
| simplify_gpt | 9m 32s | $5.55 | 0 |
| verify | 3m 43s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **74m 52s** | **$44.11** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Bryan Helmkamp <bryan@brynary.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Fabro <noreply@fabro.sh>
Co-authored-by: fabro <fabro@example.com>
Co-authored-by: fabro <fabro@fabro.sh>
|
||
|
|
c6356cbd77
|
feat: improve run board, thread, and MCP create flows (#347)
## Summary This branch improves several run-management surfaces that agents and users rely on: archived runs now stay visible and ordered correctly in the board view, pair-session messages appear in the stage Thread tab, and the `fabro_run_create` MCP tool accepts the workflow-string shorthand it advertises. ## Changes - Updates the web board cache invalidation and archived-column handling so archive/unarchive actions refresh both active and archived board queries and keep archived runs in a predictable column position. - Adds pair user/system message events to stage activity parsing, Thread rendering, search, details, and DNA timeline items. - Aligns `fabro_run_create` MCP runtime deserialization and `tools/list` schema so each run entry may be either a workflow string or a full create spec object. ## Test Plan - `cargo nextest run -p fabro-tool -p fabro-mcp-server` - `cargo nextest run -p fabro-cli stdio_server_initializes_and_lists_run_tools mcp_create_string_shorthand_deserializes_before_auth mcp_create_validation_errors_happen_before_auth_or_network mcp_create_and_search_manage_real_runs_with_cli_auth` - `cargo +nightly-2026-04-14 fmt --check --all` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
f5ec711a2c
|
Stage-based pairing API and fabro_run_pair MCP tool (#344)
Some checks are pending
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Rust / Format (push) Waiting to run
TypeScript / Build (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
## Summary
Run pairing previously required callers to supply an opaque
`agent_session_id` alongside a `stage_id` to start or target a pair
session. This leaked an internal runtime identifier across the public
HTTP API, generated TypeScript client, and would have bled into any MCP
tooling. This PR removes that coupling: the public pair API now
identifies targets by `StageId` alone, the server resolves the live
session internally, and a new `fabro_run_pair` MCP tool exposes the full
pair lifecycle without ever seeing session identifiers.
### What changed
**Public contract simplification** (`fabro-types`, OpenAPI, generated TS
client)
- `PairTarget` is now `{ stage_id, node_label }` — `node_id`, `visit`,
`agent_session_id`, `provider`, and `model` are removed.
- `PairStartRequest` accepts `{ stage_id }` instead of `{ target:
PairTargetSelector }`.
- `PairTargetSelector` and `PairTranscriptModel` types are deleted
entirely.
- `PairMessageRecord.target` (selector) replaced by
`PairMessageRecord.stage_id`.
- `PairTranscriptAssistantMessage.model` field removed.
- `MAX_PAIR_MESSAGE_BYTES` extracted as a public constant shared between
the server handler and the MCP tool.
**Internal session binding** (`SteeringHub`, server projection)
- `ActivePair` now carries `session_id: String` separately from the
public `PairRecord`. This preserves the stale-session protection that
previously relied on `target.agent_session_id`.
- Transcript matching changed from `(session_id AND stage_id)` to
`stage_id` within the already-scoped pair window sequence range —
simpler and sufficient.
- `active_api_targets` deactivation no longer does a per-target
`agent_session_id` check; it relies on the `active_steerable_stages`
lease already doing that guard.
**New `fabro_run_pair` MCP tool** (`fabro-mcp-server`)
- Actions: `status`, `start`, `get`, `message`, `end`, `transcript`.
- Validation happens before any network call; missing `run_id`, missing
`stage_id` for `start`, missing/invalid `pair_id` for other actions, and
overlong message text all return clean tool-level errors.
- `strum::IntoStaticStr` on `RunPairAction` enables the
`parse_pair_id_for_action` helper to embed the action name in error
messages without a `match`.
- MCP result schema and serialized results are covered by leakage
assertions confirming none of the removed fields surface.
**Tests**
- Negative leakage assertions added to pair DTO tests, event round-trip
tests, control-protocol tests, server handler tests, MCP validation
tests, and MCP schema test.
- Tool count updated from 5 → 6 in all CLI MCP integration tests.
- Steering hub test renamed:
`pair_start_rejects_non_selected_or_missing_target` →
`pair_start_rejects_missing_target` (session-mismatch rejection is now
an internal concern).
### Fabro Details
<details>
<summary>Ran 9 stages in 60m 29s for $31.78</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 2m 9s | – | 0 |
| preflight_lint | 2m 26s | – | 0 |
| implement | 33m 14s | $25.51 | 0 |
| simplify_opus | 15m 7s | $4.18 | 0 |
| simplify_gpt | 3m 46s | $2.08 | 0 |
| verify | 3m 9s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **60m 29s** | **$31.78** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
Co-authored-by: Bryan Helmkamp <bryan@brynary.com>
|
||
|
|
296fbddec9
|
feat(api): add ask fabro session endpoints (#342)
## Summary Adds the run-backed API surface needed for a real Ask Fabro sidebar: run readiness metadata, detailed session projections, session-scoped event listing/attach streaming, and turn control that exposes durable turn IDs and machine-readable failures. ## What Changed - Extended the OpenAPI contract and regenerated Rust/TypeScript clients for `Run.ask_fabro`, `SessionDetail`, `SessionTurn`, paginated run sessions, session event APIs, and optional client-supplied `turn_id` values. - Updated `fabro-types` and `fabro-store` so durable `run.session.*` events project active turn state, transcript messages, and the latest owning run event sequence. - Implemented server routing for session details, `/events`, `/attach`, turn conflict headers, typed turn failure codes, and cheap run readiness decoration across run responses. - Added browser helpers for POST turn streaming and session attach SSE parsing, plus an exported generated `sessionsApi`. ## Verification - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` - `cargo build --workspace` - `cargo test -p fabro-types run_session_turn_failed_defaults_code_for_old_events` - `cargo test -p fabro-store run_sessions::tests` - `cargo test -p fabro-api` - `cargo test -p fabro-server --features test-support --test it api::sessions` - `cargo test -p fabro-server --features test-support --test it api::runs` - `cd apps/fabro-web && bun test app/lib/session-stream.test.ts` - `cd apps/fabro-web && bun run typecheck` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 Codex (context unknown, medium reasoning) via [Codex](https://openai.com/codex/) |
||
|
|
54bc67017e
|
feat: Replace duration/elapsed fields with wall_time_ms and StageTiming (#343)
## Summary
Replaces the ambiguous `runtime_secs`, `elapsed_secs`, and `duration_ms`
timing fields on run/stage public API surfaces with explicit
`wall_time_ms` (elapsed clock time) and a `StageTiming` value object
that also carries `inference_time_ms`, `tool_time_ms`, and
`active_time_ms`.
This is a greenfield breaking change — no compatibility shims are
preserved.
### What changed
**API shape**
- `RunBillingStage.runtime_secs` → `RunBillingStage.timing: StageTiming`
- `RunBillingTotals.runtime_secs` → `RunBillingTotals.timing:
StageTiming`
- `RunSummary.timestamps.duration_ms` / `elapsed_secs` removed; a
top-level `timing: StageTiming | null` field added
- Stage list item `duration_secs` → `wall_time_ms`
**Web app (`apps/fabro-web`)**
- `run-billing.tsx`: `liveRuntimeSecs` → `liveWallTimeMs`; live ticking
now returns milliseconds and the footer total sums `wallTimeMs` across
rows
- `stage-sidebar.ts`: `duration_secs` → `wall_time_ms` for the per-stage
duration display
- `runs.ts`: `elapsed_secs` lookup replaced with `timing.wall_time_ms`
- `formatElapsedSecs` / `formatDurationSecs` call sites replaced with
`formatDurationMs`
**Lockfile / tooling**
- `@openapitools/openapi-generator-cli@2.20.2` added as a dev dependency
to `@qltysh/fabro-api-client` to support regenerating the TypeScript
client after schema edits; several transitive deps pulled in alongside
it.
### Design notes
- **Units are now consistent**: every timing value on run/stage surfaces
is in milliseconds; the old API mixed seconds (`runtime_secs`,
`elapsed_secs`) with milliseconds (`duration_ms`).
- **Live ticking** still works correctly: the in-flight billing row
computes `now - startedAt` in ms and sums across rows for the footer,
avoiding a server round-trip during a running stage.
- **`StageTiming.active_time_ms = inference_time_ms + tool_time_ms`** —
parallel work is summed, so run active time can exceed wall time.
- Subsystem-internal `duration_ms` fields (sandbox setup, devcontainer
lifecycle, hooks) are intentionally left unchanged; only public
run/stage timing surfaces are affected.
### Fabro Details
<details>
<summary>Ran 9 stages in 115m 53s for $108.50</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 2m 6s | – | 0 |
| preflight_lint | 2m 18s | – | 0 |
| implement | 80m 53s | $101.97 | 0 |
| simplify_opus | 21m 35s | $4.09 | 0 |
| simplify_gpt | 5m 1s | $2.44 | 0 |
| verify | 3m 11s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **115m 53s** | **$108.50** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
Co-authored-by: Bryan Helmkamp <bryan@brynary.com>
|
||
|
|
fb2174c7d0
|
feat(agent): expose Fabro run tools in sessions (#339)
## Summary API-backed Fabro agent sessions can now use the same five run-control tools that were previously MCP-only, while the server uses a single scoped `FABRO_WORKER_TOKEN` path for worker authorization. This lets server-dispatched API agents create, gather, inspect, and interact with runs without reintroducing a separate delegated run-agent token or leaking credentials into sandboxed child environments. ## Changes - Moved the reusable run-tool implementation into the new `fabro-tool` crate so MCP and API agent backends share the same tool schemas and client behavior. - Registered the five `fabro_run_*` tools for API-mode agents, with ACP sessions continuing to omit those tools. - Replaced `FABRO_RUN_AGENT_TOKEN` with scoped worker-token auth: base worker tokens keep same-run access, and `run:worker agent:run_tools` tokens can call the run-control API across runs. - Added server auth guards for run-tool actors and run-scoped-or-run-tools routes, then applied them only to the routes used by the run-tool client backend. - Kept `FABRO_WORKER_TOKEN` scrubbed from sandbox commands, hooks, ACP subprocesses, MCP env providers, and other child tool environments. ## Testing - `cargo nextest run -p fabro-server worker_token principal_middleware spawn_env --no-fail-fast` - `cargo nextest run -p fabro-cli runner --no-fail-fast` - `cargo nextest run -p fabro-workflow agent_run --no-fail-fast` - `cargo nextest run -p fabro-static --no-fail-fast` - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy -p fabro-server -p fabro-cli -p fabro-static --all-targets -- -D warnings` - `git diff --check` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) |
||
|
|
2b168b4588
|
Compute LLM cost on-read for in-flight billing stages (#345)
## Summary
The run billing page showed `—` for dollar cost on any active
(in-flight) stage because `total_usd_micros` is only computed at stage
completion. This PR prices stages whose cost is `None` at read time,
using the model and token counts already present in the projection.
## What changed
**`fabro-model/src/billing.rs`** gains two new methods on existing
types:
- `BilledTokenCounts::token_counts()` — extracts the five disjoint token
buckets, dropping the derived sum and optional cost field.
- `Catalog::price_tokens(model, tokens)` — mirrors the cost computation
from `billed_model_usage_from_llm` but callable outside the completion
path. Returns `None` for unknown models or providers with no billing
policy.
**`fabro-workflow/src/billing_rollup.rs`** —
`billing_rollup_from_projection` gains an `Option<&Catalog>` parameter.
A new private helper `stage_usage_with_cost` fills in `total_usd_micros`
on-the-fly for any stage where it is `None` and both a catalog and model
are available. All four accumulator call sites (`is_zero` check,
per-stage row, `totals`, and `by_model`) use the priced copy, keeping
the page internally consistent.
**Call sites** — the read handler (`handler/billing.rs`) passes
`Some(&catalog)` so active stages get priced. The aggregate-billing
sites in `server.rs` and the four finalization sites in `finalize.rs`
pass `None` — completed stages are already priced, and we deliberately
exclude running estimates from org-wide totals to avoid double-counting
when a run later finalizes.
## Design decisions
- **Price any stage with `total_usd_micros == None`**, not just
explicitly in-flight ones. Stages with no billing policy return `None`
again — harmless.
- **No "estimated" label** — cost-so-far is exact for tokens consumed so
far, consistent with the already-unlabeled live token count and runtime.
- In-flight **prompt** stages still show `—` because `stage.model` isn't
set until `PromptCompleted`. This is acceptable; the reported bug
concerns agent stages.
### Fabro Details
<details>
<summary>Ran 9 stages in 44m 13s for $7.17</summary>
| Stage | Duration | Cost | Retries |
|---|---|---|---|
| start | 0s | – | 0 |
| toolchain | 1s | – | 0 |
| preflight_compile | 4m 6s | – | 0 |
| preflight_lint | 4m 58s | – | 0 |
| implement | 13m 24s | $3.84 | 0 |
| simplify_opus | 10m 34s | $1.69 | 0 |
| simplify_gpt | 5m 58s | $1.64 | 0 |
| verify | 3m 40s | – | 0 |
| fmt | 3s | – | 0 |
| **Total** | **44m 13s** | **$7.17** | **0** |
</details>
<details>
<summary>Ran <code>ImplementPlan.fabro</code> (12 nodes and 15
edges)</summary>
```dot
digraph ImplementPlan {
graph [
goal="Implement and simplify",
model_stylesheet="
* { model: claude-opus-4-7; }
"
]
rankdir=LR
start [shape=Mdiamond, label="Start"]
exit [shape=Msquare, label="Exit"]
toolchain [label="Toolchain", shape=parallelogram, script="command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", max_retries=0]
preflight_compile [label="Preflight Compile", shape=parallelogram, script="cargo check -q --workspace 2>&1", max_retries=0]
preflight_lint [label="Preflight Lint", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", max_retries=0]
fix_lints [label="Fix Lints", prompt="The preflight lint step failed. Read the build output from context and fix all clippy lint warnings.", max_visits=3]
implement [label="Implement", prompt="Read the plan file referenced in the goal and implement every step. Make all the code changes described in the plan. Use red/green TDD."]
simplify_opus [label="Simplify (Opus)", prompt="@prompts/simplify.md"]
simplify_gpt [label="Simplify (GPT-55)", prompt="@prompts/simplify.md", model="gpt-55"]
verify [label="Verify", shape=parallelogram, script="cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --cargo-quiet --workspace --status-level fail 2>&1 && cargo dev docs refresh 2>&1 && cargo dev docs check 2>&1", goal_gate=true, retry_target="fixup"]
fixup [label="Fixup", prompt="The verify step failed. Read the build output from context and fix all clippy lint warnings, test failures, and generated docs errors.", max_visits=3]
fmt [label="Format", shape=parallelogram, script="cargo +nightly-2026-04-14 fmt --all 2>&1", max_retries=0]
start -> toolchain
toolchain -> preflight_compile [condition="outcome=succeeded"]
toolchain -> exit
preflight_compile -> preflight_lint [condition="outcome=succeeded"]
preflight_compile -> exit
preflight_lint -> implement [condition="outcome=succeeded"]
preflight_lint -> fix_lints
fix_lints -> preflight_lint
implement -> simplify_opus -> simplify_gpt -> verify
verify -> fmt [condition="outcome=succeeded"]
verify -> fixup
fixup -> verify
fmt -> exit
}
```
</details>
⚒️ Generated with [Fabro](https://fabro.sh)
---------
Co-authored-by: Fabro <noreply@fabro.sh>
|
||
|
|
52da187e65
|
fix(workflow): follow symlinked artifact roots (#338)
Fixes https://github.com/fabro-sh/fabro/issues/335 ## Summary - Run artifact discovery with `find -H` so a symlinked sandbox working-directory root is traversed. - Keep existing behavior for symlinks discovered inside the tree by preserving the current `-not -type l -type f` filter. - Add workflow integration coverage for artifact collection when the local sandbox working directory itself is a symlink. ## Test plan - `cargo nextest run -p fabro-workflow asset_collection_local_sandbox_symlink_working_directory` - `cargo nextest run -p fabro-workflow artifact_snapshot` - `cargo nextest run -p fabro-workflow asset_collection_local_sandbox` - `cargo +nightly-2026-04-14 fmt --check --all` |
||
|
|
02fabb7b47
|
fix(workflow): capture configured artifacts once (#337)
## Summary Fixes artifact promotion for configured `[run.artifacts].include` globs by collecting matching files after each stage regardless of mtime and surfacing failed discovery commands as collection failures. The workflow lifecycle now keeps a per-run ledger keyed by `(path, content_sha256)`, rebuilt from existing `artifact.captured` events, so unchanged files are persisted and emitted once while changed content at the same path can still be captured again. Fixes fabro-sh/fabro#335. ## Tests - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo nextest run -p fabro-workflow artifact_snapshot` - `cargo nextest run -p fabro-cli unchanged_matching_artifact_is_captured_once_across_stages` - `cargo nextest run -p fabro-cli acp_artifacts_are_listed_when_touched_file_mtime_precedes_attempt_start` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) --------- Co-authored-by: Jess Martin <27258+jessmartin@users.noreply.github.com> |
||
|
|
15ff38fa53
|
feat(template): resolve template error locations (#333)
## Summary Fixes `fabro-sh/fabro#330` by making template partials that reference missing inputs validate structurally with a warning instead of failing the validate command. The template crate now owns MiniJinja semantic error classification and source-location mapping, so workflow diagnostics can consume already-resolved template locations instead of remapping fragment spans itself. ## What Changed - Added `TemplateErrorLocation` and `TemplateSourceOrigin` APIs to report source name, line, column, and span from `fabro-template`. - Classified wrapped MiniJinja errors by their deepest semantic cause, preserving the original source chain for renderer context. - Added fragment-origin rendering paths so attribute fragments embedded in full workflow source report locations in the original source text. - Removed workflow-side source span remapping from template diagnostics; workflow now only adds owner, node/edge, severity, rule, and fix context. - Added regression coverage for include/import/from/extends undefined variables and the CLI `fabro validate` partial fixture. ## Test Plan - `cargo nextest run -p fabro-template` - `cargo nextest run -p fabro-workflow transforms::variable_expansion transforms::file_inlining` - `cargo nextest run -p fabro-cli --test it cmd::validate` - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy -p fabro-template -p fabro-workflow -p fabro-cli --all-targets -- -D warnings` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 (context not reported, reasoning not reported) via [Codex](https://openai.com/codex) |
||
|
|
bf96baa9f0
|
feat(server): add run pairing API (#312)
## Summary
Adds the server-side run pairing surface for joining one active API-mode
agent session, sending pair messages, reading a compact transcript, and
ending pairing explicitly before workflow release continues.
This PR wires the feature end to end:
- adds OpenAPI paths and shared `fabro-types` DTOs for pair lifecycle,
messages, transcript entries, and run event details
- adds typed `RunEvent` variants for pair lifecycle and pair-scoped
user/system messages
- extends the workflow steering hub and agent session drain path with
typed pair control items, single-target validation, pair parking, and
pair end/resume behavior
- extends worker JSONL control and server transports for pair
start/message/end while preserving existing
steer/interrupt/answer/cancel behavior
- adds Axum handlers for `/api/v1/runs/{id}/pair`, pair messages, pair
transcript, and `/api/v1/runs/{id}/events/{seq}`
- adds `fabro-client` helpers for the new endpoints
## Notes
The subprocess path does not add a bidirectional worker ack channel in
this PR. Instead, the HTTP pair handlers only return lifecycle/message
success after the corresponding durable runtime event is observed, so
mpsc enqueue success alone is not treated as API success.
The plan checklist in
`docs/superpowers/plans/2026-05-18-server-side-run-pairing-api-events.md`
is included with that distinction left visible.
## Verification
- `cargo build -p fabro-api`
- `cargo check -p fabro-api -p fabro-client -p fabro-agent -p
fabro-workflow -p fabro-interview -p fabro-server`
- `cargo +nightly-2026-04-14 fmt --check --all`
- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D
warnings`
- `cargo nextest run -p fabro-api pair
run_event_round_trips_pair_lifecycle_events
run_event_round_trips_agent_pair_messages`
- `cargo nextest run -p fabro-workflow pair`
- `cargo nextest run -p fabro-interview pair`
- `cargo nextest run -p fabro-server pair
subprocess_answer_transport_pair_commands_enqueue_control_messages steer
interrupt`
|
||
|
|
5bfd115339
|
[codex] Add ACP steering support (#329)
## Summary - Adds a backend-neutral live control abstraction so steering, interrupt, and interrupt+steer no longer depend on API-only session handles. - Reworks ACP sessions into a live protocol loop that uses ACP `session/prompt` for follow-up steers and ACP `session/cancel` for interrupts without restarting the process. - Registers ACP sessions as steerable, removes the stale non-steerable server/UI/API path, preserves ACP projection metadata, and keeps unsupported backends out of the steerability gate. ## Validation - `LC_ALL=C cargo nextest run --workspace --no-fail-fast` (5,833 passed, 178 skipped) - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` - `cargo +nightly-2026-04-14 fmt --check --all` - `git diff --check` - `cd apps/fabro-web && LC_ALL=C ASDF_NODEJS_VERSION=20.13.1 ASDF_BUN_VERSION=1.3.11 bun test` (396 passed) - `cd apps/fabro-web && LC_ALL=C ASDF_NODEJS_VERSION=20.13.1 ASDF_BUN_VERSION=1.3.11 bun run typecheck` - `cargo build -p fabro-api` - `cd lib/packages/fabro-api-client && LC_ALL=C ASDF_BUN_VERSION=1.3.11 bun run typecheck` --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> Co-authored-by: Bryan Helmkamp <bryan@brynary.com> |
||
|
|
9bdf30ad86
|
refactor(hooks): centralize run location handling (#325)
## Summary - Fix host command hooks to use the submitter/source directory instead of a sandbox-only working directory. - Introduce `HookExecutionContext` and `RunLocations` so host source, sandbox work, and run scratch paths are explicit. - Route lifecycle hooks and tool hooks through the shared hook execution context instead of rebuilding cwd pairs at call sites. ## Test Plan - `cargo nextest run -p fabro-hooks` - `cargo check -p fabro-workflow --tests` - `cargo +nightly-2026-04-14 fmt --package fabro-hooks --package fabro-workflow --check` - `cargo +nightly-2026-04-14 clippy -p fabro-hooks -p fabro-workflow --all-targets -- -D warnings` --------- Co-authored-by: Jess Martin <jessmartin@gmail.com> |
||
|
|
9f6823b10d
|
fix(workflow): honor human gate timeout defaults (#323)
## Summary - add a regression test for unanswered human gate timeouts with `human.default_choice` - route human gate `timeout` through the interview timeout path so it emits `interview.timeout` and selects the default target - make timeout ownership explicit for handlers that consume `node.timeout()`, including command and ACP handlers Fixes #317 ## Testing - cargo nextest run -p fabro-workflow --test it human_gate_timeout_routes_to_default_choice_when_unanswered - cargo nextest run -p fabro-workflow timeout_policy built_in_handlers_that_consume_node_timeout_manage_it_themselves agent_handler_delegates_timeout_policy_to_backend - cargo nextest run -p fabro-workflow script_handler_timeout script_handler_timeout_error_includes_output_tails writes_script_timing_json_on_timeout timeout_causes_fail_status_record - cargo nextest run -p fabro-workflow wait_human - cargo +nightly-2026-04-14 fmt --check --all - cargo +nightly-2026-04-14 clippy -p fabro-workflow --all-targets -- -D warnings --------- Co-authored-by: Jess Martin <jessmartin@gmail.com> |
||
|
|
21c408d647
|
fix(workflow): allow workflow-root template partials (#322)
## Summary - Add a CLI validation regression for workflow-root prompt partials included from nested prompt files. - Allow bundled prompt templates to resolve sibling partials from the workflow root instead of jailing each prompt file to its own directory. - Refactor include handling so manifest discovery and runtime rendering share rooted template sources, include normalization, root containment checks, and FileResolver-backed TemplateStore loading. ## Testing - `cargo nextest run -p fabro-template` - `cargo nextest run -p fabro-manifest` - `cargo nextest run -p fabro-workflow` - `cargo nextest run -p fabro-cli cmd::validate` - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy -p fabro-template -p fabro-manifest -p fabro-workflow -p fabro-cli --all-targets -- -D warnings` --------- Co-authored-by: Aleksi Asikainen <1086393+salieri@users.noreply.github.com> |
||
|
|
29b7cc0de0
|
feat(workflow): enforce strict api/acp backends (#307)
## Summary
This PR makes agent execution a strict two-backend contract: API-backed
stages use Fabro-owned model/provider auth, while ACP-backed stages
launch a user-supplied stdio process that owns its own auth and tools.
That removes the legacy CLI backend and prevents ACP execution from
accidentally resolving or forwarding provider credentials.
## Changes
- Replaces the old `api`/`cli`/`acp` backend model with `AgentBackend {
api, acp }`, with `backend=\"cli\"` rejected and migrated toward
explicit ACP process configuration.
- Splits ACP process configuration into `acp.command` for shell command
strings and `acp.config` for JSON stdio configs, while rejecting legacy
`acp_command`.
- Restricts ACP to `agent` nodes and rejects API-only attributes such as
`model`, `provider`, `reasoning_effort`, `max_tokens`, and `speed` on
ACP nodes.
- Deletes the workflow CLI runtime, CLI credential resolver surface, CLI
live smoke tests, and `agent.cli.*` event handling.
- Updates ACP events and projections to report process identity
(`command`, optional `config_name`) rather than provider/model metadata.
- Updates import/stylesheet propagation, CLI workflow smoke coverage,
server steering tests, and web model extraction for the new
event/backend contract.
## Validation
- `cargo check -p fabro-auth -p fabro-acp -p fabro-workflow -p fabro-cli
--all-targets`
- `cargo nextest run -p fabro-auth -p fabro-acp -p fabro-validate -p
fabro-store -p fabro-workflow --lib`
- `cargo nextest run -p fabro-acp`
- `cargo nextest run -p fabro-cli --test it
workflow::acp::acp_backend_workflow`
- `cargo nextest run -p fabro-workflow --test it
codergen_without_backend_simulated`
- `cargo nextest run -p fabro-workflow --test it
import_e2e_through_engine`
- `cargo nextest run -p fabro-workflow --test it stylesheet_application`
- `cargo nextest run -p fabro-server
steer_with_active_acp_stage_returns_non_steerable_conflict`
- `cargo nextest run -p fabro-server
active_acp_stage_marker_clears_on_terminal_paths`
- `cargo nextest run -p fabro-types
agent_backend_accepts_only_api_and_acp`
- `cd apps/fabro-web && bun test app/routes/run-stages.test.ts`
- `cd apps/fabro-web && bun run typecheck`
- `cargo +nightly-2026-04-14 fmt --check --all`
- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D
warnings`
---
[](https://github.com/EveryInc/compound-engineering-plugin)
🤖 Generated with GPT-5 via [Codex](https://openai.com/codex)
---------
Co-authored-by: Peter Bell <4843+PeterBell@users.noreply.github.com>
|
||
|
|
cd74013d06
|
refactor(auth): split credential sources and vault schemas (#306)
## Summary Compared with `origin/main`, this PR splits credential storage and credential references into explicit types. Vault secrets now distinguish `token`, `oauth`, and `file` payloads, while runtime/model configuration points to credentials through explicit `env:<NAME>` and `vault:<NAME>` source refs. ## Changes - Replaces the old `environment`/`credential` secret schema vocabulary with `token`/`oauth`/`file` across OpenAPI, Rust API tests, generated TypeScript models, CLI/docs references, and the changelog. - Updates auth resolution, refresh, provider strategies, workflow LLM handling, server diagnostics, install flows, run manifests, and secret handlers to consume typed vault entries and explicit credential sources. - Updates provider catalog TOMLs and config parsing so provider auth and extra headers use `vault` refs instead of ambiguous `credential` refs. - Updates CLI install/login/run/secret paths and integration tests to write and read the new credential shapes. - Removes the temporary legacy vault migration and empty-vault fallback, then centralizes provider vault secret-name lookup and Codex API credential shaping. ## Verification - `cargo +nightly-2026-04-14 fmt --all` - `cargo +nightly-2026-04-14 clippy -p fabro-auth -p fabro-model -p fabro-config -p fabro-vault -p fabro-server -p fabro-cli --all-targets -- -D warnings` - `ulimit -n 4096 && cargo nextest run -p fabro-auth -p fabro-model -p fabro-config -p fabro-vault -p fabro-server -p fabro-cli` (`1938` passed, `35` skipped) |
||
|
|
6a86ced77c
|
fix(server): persist manifest metadata names (#302)
When Fabro creates a detached run through the server, the resulting run metadata should still read like something a human can trust at a glance. Before this change, those runs could persist with `settings.project.name` and `settings.workflow.name` left `null` even though Fabro already had enough local context to infer them. That made `inspect` output look half-populated and made it harder to tell whether the saved run state was complete. This fixes that trust gap in the server-backed manifest flow. ## Summary - backfill missing manifest-backed project and workflow names during server run preparation - prefer explicit `[workflow].name` from bundled `workflow.toml`, then fall back to graph name or workflow slug - cover both manifest preparation and persisted run-state behavior with server tests ## Testing - cargo test -p fabro-server prepare_manifest_backfills_missing_project_and_workflow_names -- --nocapture - cargo test -p fabro-server prepare_manifest_preserves_explicit_project_and_workflow_names -- --nocapture - cargo test -p fabro-server create_run_persists_backfilled_project_and_workflow_names -- --nocapture --------- Co-authored-by: Bryan Helmkamp <bryan@brynary.com> |
||
|
|
492aba7fff
|
fix: MiniJinja can't find partials (#301)
Some checks are pending
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
TypeScript / Build (push) Waiting to run
Fixes an issue where use of MiniJinja [`include`](https://jinja.palletsprojects.com/en/stable/templates/#include) control structure (`{% include "filename.ext" %}`) causes a render error `template not found: tried to include non-existing template "filename.ext"` ### Example broken diagram ``` dot digraph ValidatePlan { start [shape=Mdiamond, label="Start"] exit [shape=Msquare, label="Exit"] test_inline_prompt [label="moo" prompt="{% include 'test.tpl.md' %}"] // ^^^^^^^^^^^^^^^^^^^^^^^^^ start -> test_inline_prompt -> exit } ``` ### Fix The core issue was that template rendering knew the source name for diagnostics, but did not have a loader rooted at the prompt/goal file location. Includes therefore failed even when the included file existed next to the rendered file. The fix adds optional loader support to `fabro-template`, then wires workflow rendering to the existing `FileResolver` so includes resolve relative to the file currently being rendered. For `fabro validate`, there was a second manifest-specific problem: validation runs through a bundled manifest, and the manifest builder only bundled explicit `prompt.md` / `goal.md` files, not static MiniJinja `include` dependencies inside those files. The manifest builder now scans prompt/goal template text for literal `{% include "file" %}` / `{% include 'file' %}` references and bundles those files too. Missing or unsafe include names are left for MiniJinja/runtime validation rather than expanding scope. (For clarity: The fix does not support variables or arrays in `include`.) |
||
|
|
302e2445b4
|
refactor(model): move provider facts into catalog (#298)
## Summary Moves provider-specific facts out of `AdapterKind` metadata and into provider catalog data, leaving adapters responsible for runtime protocol behavior. This makes providers that share an adapter mostly TOML-driven while still surfacing adapter construction failures during readiness checks. ## What Changed - Provider TOML now owns auth mode, API-key/header policy, billing policy, agent profile, base URLs/env overrides, extra headers, and probe markers. - Auth, install, config, diagnostics, and server flows resolve provider credentials from catalog auth config, including API-key, header-only, and no-auth providers. - LLM client registration now reports adapter construction failures, validates final adapter requests before HTTP dispatch, and preserves custom primary auth headers. - Billing and docs now use provider-owned billing policy instead of adapter metadata, and the old adapter metadata surface is removed. ## Reviewer Notes OpenAI-compatible `base_url` validation now happens during adapter/client registration rather than catalog build. That keeps catalog parsing adapter-agnostic while still letting readiness and model listing reflect providers that cannot register. ## Verification - `cargo check -p fabro-model -p fabro-auth -p fabro-llm -p fabro-server -p fabro-cli` - `cargo nextest run -p fabro-llm -- adapter_registry` - `cargo nextest run -p fabro-model -- catalog` - `cargo nextest run -p fabro-auth -- api_key` - `cargo nextest run -p fabro-server -- install` - `cargo +nightly-2026-04-14 fmt --check --all` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
2ba04be181
|
feat(template): add source-aware diagnostics (#292)
## Summary Template failures from `fabro run` and structural warnings from `fabro validate` now preserve source provenance through rendering, workflow transforms, API serialization, and CLI display. Diagnostics can point at the actual workflow, import, or prompt file with node/attribute context instead of surfacing MiniJinja's generic `<string>` source. ## What Changed - Added named MiniJinja render APIs plus miette-aware `TemplateError` metadata for source names, source text, spans, and labels. - Reworked workflow template expansion so inline attributes, imported workflows, and `@prompt` files render with file and owner context. - Split strict run behavior from structural validate behavior: run-start still hard-fails on missing inputs, while validate emits source-aware warnings and continues linting. - Extended validation diagnostics through Rust structs, OpenAPI, server DTO mapping, and CLI rendering with optional source path, line, column, span, and related metadata. - Added regression coverage across template rendering, workflow transforms, CLI output, and the server validate endpoint. ## Verification - `cargo nextest run -p fabro-template` - `ulimit -n 4096 && cargo nextest run -p fabro-workflow --no-fail-fast` - `cargo nextest run -p fabro-cli bare_fabro_with_unbound_inputs_validates_structurally_with_warning run_rejects_unbound_template_inputs_before_creating_remote_run` - `cargo nextest run -p fabro-server validate_endpoint_returns_template_source_coordinates` - `cargo build -p fabro-api` - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) --------- Co-authored-by: Aleksi Asikainen <1086393+salieri@users.noreply.github.com> |
||
|
|
b9497517a5
|
feat(llm): support agent profile overrides (#291)
## Summary Adds catalog-level `agent_profile` overrides so custom providers and individual models can choose Anthropic, OpenAI, or Gemini agent behavior independently from their adapter default. The effective precedence is model override, then provider override, then adapter metadata. ## What Changed - Added typed provider/model `agent_profile` settings in `fabro-config` and `fabro-model`, with serde/strum support for `anthropic`, `openai`, and `gemini`. - Centralized effective profile resolution in the catalog, including provider alias canonicalization and a guard against unrelated model overrides leaking across providers. - Updated run startup, API sessions, CLI/ACP backends, prompt project-memory discovery, and standalone agent startup to use the resolved catalog profile. - Documented provider-level and model-level `agent_profile` configuration in the public model and user configuration docs. No OpenAPI or model-list response shape changes are included. ## Validation - `cargo nextest run -p fabro-model -p fabro-config -p fabro-workflow -p fabro-agent` passed: 1826 passed, 125 skipped. - `cargo +nightly-2026-04-14 fmt --check --all` passed. - `cargo +nightly-2026-04-14 clippy -p fabro-model -p fabro-config -p fabro-workflow -p fabro-agent --all-targets -- -D warnings` passed. - `git diff --check` passed. - `cargo nextest list -p fabro-dev` confirmed there is no docs-options reference test target to run. ## Post-Deploy Monitoring & Validation Watch workflow and agent-session logs for provider/model resolution errors, unexpected project-memory file selection, or CLI/ACP launch command mismatches on custom catalog providers. Healthy signal: custom provider/model runs start normally and use the intended profile-specific behavior. Failure trigger: repeated `Provider ... is not configured` errors, missing expected project memory, or profile-specific agent startup failures after configuring `agent_profile`. Mitigation is to remove the override from config or revert this PR. Validation window: first deploy cycle after merge; owner: release/on-call engineer. --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) |
||
|
|
2b53917759
|
fix(validate): treat undefined template vars in @file prompts as warnings (#290)
## Summary
`fabro validate` had inconsistent behavior for undefined template
variables depending on whether the prompt was inline or loaded via an
`@file` reference. Inline `{{ inputs.foo }}` produced a warning and
validation passed; the same expression inside a `@file`-imported prompt
produced a hard validation error.
Fixes #286.
## Root cause
Two template-rendering passes with different strictness, applied to
disjoint inputs:
1. **DOT-source pass**
(`lib/crates/fabro-workflow/src/operations/create.rs`) honored
`RenderMode::Structural` for `fabro validate` — undefined variables
downgraded to a `Severity::Warning` diagnostic, then lenient render
finished the job.
2. **Per-attribute pass**
(`lib/crates/fabro-workflow/src/transforms/variable_expansion.rs`)
inside `TemplateTransform` was always strict and had no `RenderMode`
awareness. Because `FileInliningTransform` runs *before*
`TemplateTransform`, expressions inside `@file` content only ever
encountered the strict pass.
## Fix
- Plumb `RenderMode` through `TransformOptions` into
`TemplateTransform`.
- In `RenderMode::Structural`, the transform catches
`TemplateError::UndefinedVariable` per attribute, emits a warning
diagnostic, and falls back to `render_lenient`.
- Diagnostics flow through a new `Transformed.diagnostics` field into
`Validated` alongside lint output.
- Diagnostics now include `node_id` when the undefined variable was
found inside a node attribute, which is more useful than the previous
"at line 1" location.
- `RenderMode` and the shared `template_undefined_variable_diagnostic`
helper moved to `pipeline/types.rs` so the transform layer can reach
them without a circular dep.
Strict mode (`fabro run`, preflight) is unchanged — undefined inputs
still hard-fail before a run is created.
## Behavior
Illustrative output shapes (variable names and line numbers depend on
the fixture):
Inline prompt (unchanged):
```
warning: undefined template variable `inputs.<name>` at line <n> (template_undefined_variable)
Validation: OK
```
`@file`-imported prompt (previously a hard error, now matches inline —
node-attributed instead of line-attributed):
```
warning [node: <id>]: undefined template variable `inputs.<name>` in node `<id>` (template_undefined_variable)
Validation: OK
```
## Test plan
- [x] `cargo nextest run --workspace` — 5773/5773 passing
- [x] `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D
warnings` clean
- [x] `cargo +nightly-2026-04-14 fmt --check --all` clean
- [x] New regression test
`bare_fabro_with_unbound_inputs_in_imported_prompt_validates_structurally_with_warning`
in `lib/crates/fabro-cli/tests/it/cmd/validate.rs` against new fixture
`test/templated_unbound_imported/`
- [x] Existing
`bare_fabro_with_unbound_inputs_validates_structurally_with_warning` and
`strict_render_hard_fails_on_unbound_inputs` still pass — verifies
inline structural and run-start strict behavior are both preserved
- [x] Manual reproduction of the exact inputs from the issue now
succeeds with a warning
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Aleksi Asikainen <1086393+salieri@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
56c7f627c4
|
feat(session): add server-backed agent sessions (#278)
## Summary
Adds the first server-backed Fabro agent session slice: persistent
session records, durable turn/event storage, HTTP session APIs, SSE turn
streaming, generated clients, and a new `fabro session -p <prompt>` CLI
path.
## What Changed
- Adds shared session IDs, records, statuses, event envelopes, and
message DTOs in `fabro-types`, with OpenAPI replacements in `fabro-api`.
- Renames the agent runtime transcript item from `Turn` to `Message` and
adds conversion between runtime history and persisted `SessionMessage`
records.
- Introduces a file-backed `SessionStore` for session metadata, turns,
full transcripts, and append-only events under local storage.
- Wires server session routes for create/list/read/update/delete, turn
submission, event replay, interrupt requests, and session-scoped tools.
- Implements streamed turn execution with durable events persisted
before SSE broadcast, active-turn conflict handling, local same-machine
`working_dir` validation, and noninteractive permission denials.
- Adds `fabro-client` helpers and the `fabro session -p` command, plus
regenerated TypeScript API client files.
## Notes
V1 intentionally keeps session execution local to same-machine server
targets. Remote clone-backed session sandboxes, interactive REPL/TUI
behavior, warm session pooling, and real tool discovery for
`/sessions/{id}/tools` remain follow-up work.
## Verification
- `cargo build --workspace`
- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D
warnings`
- `cargo +nightly-2026-04-14 fmt --check --all`
- `cargo test -p fabro-store session_store_contract_tests --lib`
- `cargo test -p fabro-agent
history::tests::session_message_roundtrip_preserves_runtime_history
--lib`
- `cargo test -p fabro-server 'session_' --lib`
- `cargo test -p fabro-server --features test-support --test it
openapi_conformance -- --nocapture`
- `cargo test -p fabro-cli --test it cmd::session:: -- --nocapture`
- `cd lib/packages/fabro-api-client && bun run typecheck`
- `git diff --check`
---
[](https://github.com/EveryInc/compound-engineering-plugin)
🤖 Generated with GPT-5 via [Codex](https://openai.com/codex)
|
||
|
|
4f5e3b78f8
|
refactor: remove compatibility shims (#281)
## Summary Simplifies the greenfield PR/run schema surface by collapsing alias-only type shims and removing legacy compatibility paths that kept old wire shapes and workflow names alive. ## Changes - Use canonical `Run`, `PullRequestLink`, `PullRequestResponse`, `BoardColumn`, `WorkflowSettings`, SWR `Key`, and `SteerRunRequest` names directly across Rust and web code. - Remove legacy PR/event deserialization compatibility for old PR records and command output fields, with tests updated to reject stale wire shapes. - Drop obsolete workflow aliases for `agent_loop`, `one_shot`, `codergen_mode`, and `stack.child_dotfile`, then update docs and tests to the current names. ## Verification - `git diff --check` - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` - `cargo nextest run -p fabro-types -p fabro-api -p fabro-client -p fabro-store -p fabro-server -p fabro-workflow -p fabro-cli` - `cd apps/fabro-web && bun run typecheck` - `cd apps/fabro-web && bun test` --- [](https://github.com/EveryInc/compound-engineering-plugin) Generated with GPT-5 via [Codex](https://openai.com/codex) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
2296ba6ea8
|
feat(runs): add parent run links (#271)
## Summary
Adds orchestration-only parent links between runs without merging them
into fork or rewind lineage. Runs can now be created under a parent,
linked to a different parent, or unlinked through event-sourced
mutations that rebuild summaries and projections from the run event
stream.
## Changes
- Adds optional `parent_id` to run manifests, public run summaries, run
projections, `run.created`, OpenAPI, and the generated TypeScript API
client.
- Adds `PUT /api/v1/runs/{id}/parent` and `DELETE
/api/v1/runs/{id}/parent` for mutable parent links across any run state,
including terminal or archived runs.
- Records `run.parent.linked` and `run.parent.unlinked` events with
actor metadata and previous/current parent IDs.
- Validates parent changes in the API path: parent must exist for new
links, self-parenting is rejected, cycles are rejected, and same-parent
or already-root operations are idempotent no-ops.
- Adds `parent_id` filtering to run listing while preserving dangling
historical parent references after parent deletion.
## Validation
- `cargo +nightly-2026-04-14 fmt --check --all`
- `cargo build -p fabro-api`
- `cargo check -p fabro-workflow -p fabro-store -p fabro-server`
- `cargo nextest run -p fabro-types -p fabro-store`
- `cargo nextest run -p fabro-server --features test-support
create_run_can_set_parent_and_list_children
link_relink_and_unlink_parent_are_idempotent
parent_link_validation_rejects_missing_self_and_cycles
deleting_parent_leaves_child_parent_id_as_historical_reference`
- `cargo nextest run -p fabro-api
run_summary_json_matches_openapi_shape`
- `cd lib/packages/fabro-api-client && bun run typecheck`
Known unrelated broad-suite blocker: `cargo nextest run -p fabro-server
--features test-support get_graph_returns_svg` currently returns 500
because the render subprocess emits test-harness output instead of SVG.
---
[](https://github.com/EveryInc/compound-engineering-plugin)
🤖 Generated with GPT-5 (unknown context, medium reasoning) via
[Codex](https://openai.com/codex)
|
||
|
|
ae55bded81
|
fix(sandbox): clone Daytona repos under /home/daytona/repos (#285)
## Summary
Daytona's default snapshot runs as the `daytona` user (uid 1001), which
lacks write permission on `/`. With `run.clone.enabled = true`, sandbox
init failed at `fs.create_folder("/repos", ...)` with HTTP 400, before
the first workflow stage could run:
```
sandbox.git.failed error="Failed to create Daytona repos root" causes=["HTTP 400"]
run.failed
```
Root cause: the Daytona provider was using Docker's root-level `/repos`
layout. Docker works because its containers run as root; Daytona's
default sandbox user does not.
**Fix:** move `REPOS_ROOT` for Daytona to `/home/daytona/repos`,
alongside the existing `/home/daytona/workspace`. The path is writable
by the default sandbox user, the symlink layout is unchanged
(`/home/daytona/workspace/<repo>` →
`/home/daytona/repos/<owner>/<repo>`),
and Docker keeps its existing `/repos` path.
**Bonus — better error diagnostics.** A new `wrap_fs_error(operation,
path, error)` helper in the Daytona provider:
- includes the attempted path in the message (was just "Failed to create
Daytona repos root" with no indication of which path);
- classifies HTTP 400 as a likely permission issue and points at
snapshot configuration;
- classifies HTTP 401/403 as an API key permissions issue;
- preserves the underlying `DaytonaError` in the source chain
(per `docs/internal/error-handling-strategy.md` — verified by walking
`Error::source()` in the regression test).
So if this class of failure recurs (custom snapshot, future path
changes, ...) the user gets:
> Failed to create Daytona repos root '/home/daytona/repos' failed
> (HTTP 400). This usually means the sandbox user lacks write permission
> on the parent directory. If you're using a custom Daytona snapshot,
> ensure the sandbox user can write to '/home/daytona/repos', or use a
> path under the user's home directory (e.g. /home/daytona/...).
instead of:
> Failed to create Daytona repos root
> HTTP 400
## Test plan
- [x] `cargo build --workspace`
- [x] `cargo nextest run -p fabro-sandbox --features daytona` — 142/142
pass
- [x] `cargo nextest run -p fabro-types -p fabro-workflow` — 1365/1365
pass
- [x] New unit test `wrap_fs_error_classifies_http_400_and_403` —
asserts
top-level message contains path + hint AND walks the source chain
to prove `DaytonaError::Api { status_code: 400, .. }` is preserved
- [x] `cargo +nightly-2026-04-14 fmt --check --all`
- [x] `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D
warnings`
- [x] **Live regression**: `daytona_clone_layout_live_smoke` against the
default `daytona-medium` snapshot — failed with `Failed to create
Daytona repos root / HTTP 400` before the change; passes
end-to-end after (provisions sandbox → clones repo → verifies
symlink + HEAD match in 2.5s)
## Related
- Closes #284 (thanks @jessmartin for the report, diagnosis, and
proposed fix)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Jess Martin <27258+jessmartin@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
66519ee12a
|
feat(errors): add structured failure diagnostics (#277)
## Summary
- Make `FailureDetail` the canonical rich diagnostic shape for stage and
terminal failures, with terminal `RunFailure` carrying `{ reason, detail
}`.
- Preserve cause chains and move process stdout/stderr diagnostics into
sanitized `exec_output_tail` instead of embedding them in messages or
causes.
- Update ACP error plumbing, CLI/server/store rendering, OpenAPI, and
the generated TypeScript API client for the nested failure contract.
Closes #273
## Test Plan
- `cargo nextest run -p fabro-types -p fabro-core -p fabro-acp -p
fabro-api -p fabro-store -p fabro-server -p fabro-workflow -p fabro-cli
--no-fail-fast -E 'not test(/returns_svg/)' --status-level fail
--final-status-level fail`
- `cargo +nightly-2026-04-14 fmt --check --all`
- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D
warnings`
- `bun run typecheck` in `lib/packages/fabro-api-client`
- `bun run typecheck` in `apps/fabro-web`
|
||
|
|
87950295bd
|
refactor(llm): split provider identity from adapters (#280)
## Summary This PR separates provider identity from adapter behavior across the LLM stack. Provider IDs now represent catalog rows and provider metadata, while adapter/profile routing owns protocol behavior for Anthropic, OpenAI, Gemini, and OpenAI-compatible providers. ## Changes - Replace the shared `fabro_model::Provider` enum with open-ended `ProviderId` catalog identity and typed `AdapterKind` metadata. - Route auth, CLI, ACP, workflow, memory selection, profile construction, and LLM client registration through catalog provider rows instead of provider-ID fallbacks. - Move API-key URL/header/env metadata into provider catalog/auth flows and require configured provider rows for credential-backed clients. - Simplify billing to `algorithm`-tagged OpenAI, Anthropic, and Gemini shapes; OpenAI-compatible adapters bill through the OpenAI algorithm. - Remove greenfield compatibility paths for old provider aliases, legacy provider-tagged billing JSON, and the `openai_compatible` pseudo-provider env fallback. - Update fixtures and tests to exercise catalog-driven Kimi/Zai/Minimax/Inception/custom OpenAI-compatible routing. ## Validation - `cargo test --no-run -p fabro-model -p fabro-auth -p fabro-agent -p fabro-workflow -p fabro-server -p fabro-llm -p fabro-api -p fabro-cli -p fabro-store -p fabro-static` - `cargo nextest run -p fabro-model -p fabro-auth -p fabro-agent -p fabro-workflow -p fabro-server --no-fail-fast` - `cargo nextest run -p fabro-llm -p fabro-api -p fabro-cli -p fabro-store -p fabro-static --no-fail-fast` - `cargo +nightly-2026-04-14 fmt --check --all` - `git diff --check` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) |
||
|
|
d09e6cde33
|
feat(pr): support GitHub pull request associations (#270)
## Summary
Adds event-sourced pull request association management for runs while
preserving Fabro-created PR creation. A run can now store a current
GitHub PR association, replace it by linking another GitHub PR URL, and
remove it through an unlink event.
## What Changed
- Added `pull_request.linked` and `pull_request.unlinked` events,
projection replay support, and optional PR metadata fields in shared
pull request records.
- Added API, server, and client support for `PUT
/runs/{id}/pull_request` and `DELETE /runs/{id}/pull_request`; linking
accepts GitHub PR URLs, infers owner/repo/number, and captures live
GitHub title and branch metadata when available.
- Added `fabro pr link` and `fabro pr unlink`, updated `fabro pr view`,
and kept create/merge/close behavior guarded to GitHub PRs with usable
coordinates.
- Updated web UI rendering and internal event docs so stored PR links
display cleanly when live GitHub details are unavailable.
## Testing
- `cargo +nightly-2026-04-14 fmt --check --all`
- `git diff --check`
- `cargo build -p fabro-api`
- `cargo nextest run -p fabro-types -p fabro-store -p fabro-server -p
fabro-cli`
- `bun run typecheck` in `lib/packages/fabro-api-client`
- `bun run typecheck` in `apps/fabro-web`
- `bun test` in `apps/fabro-web`
Refs https://github.com/fabro-sh/fabro/issues/235
---
[](https://github.com/EveryInc/compound-engineering-plugin)
🤖 Generated with GPT-5 via [Codex](https://openai.com/codex)
---------
Co-authored-by: Haroldo Olivieri <6575718+haroldolivieri@users.noreply.github.com>
|
||
|
|
b0847180b6
|
Fix custom provider resolution in exec and model list (#276)
## Summary - allow `fabro exec` to use configured custom provider IDs from the resolved LLM catalog - route direct exec sessions through the same catalog-aware provider/profile resolution used by workflow runs - stop `fabro-client::list_models` from rejecting non-built-in provider filters client-side - update CLI snapshots and add regression tests for custom-provider exec and model listing ## Repro With a configured provider like: ```toml [llm.providers.bedrock] adapter = "openai_compatible" base_url = "https://.../v1" [cli.exec.model] provider = "bedrock" name = "bedrock-claude-sonnet-4-6" ``` these paths diverged: - `fabro run ... --model bedrock-claude-sonnet-4-6` worked - `fabro model list` showed `bedrock-*` models - `fabro exec "..."` failed with `unknown provider: bedrock` - `fabro model test --provider bedrock` failed with the same client-side error ## Root cause There were two separate built-in-only assumptions: 1. `fabro-agent` direct CLI paths parsed provider strings into the built-in `Provider` enum and built a default catalog, so configured provider IDs from `settings.toml` were invisible. 2. `fabro-client::list_models()` parsed the optional provider filter into the same built-in enum before calling the server, so custom provider filters never reached the API. ## Validation - `cargo check -p fabro-cli -p fabro-agent -p fabro-client` - `cargo test -p fabro-agent resolve_provider_accepts_custom_catalog_provider -- --nocapture` - `cargo test -p fabro-client list_models_allows_custom_provider_filters -- --nocapture` - `cargo test -p fabro-cli exec_accepts_configured_custom_provider_from_settings -- --nocapture` - `cargo test -p fabro-cli list_invalid_provider_errors -- --nocapture` - `cargo test -p fabro-cli help -- --nocapture` --------- Co-authored-by: Bryan Helmkamp <bryan@brynary.com> |
||
|
|
78941a2b84
|
test(model): prepare catalog-only providers (#264)
## Summary Prepares the model catalog tests for catalog-only built-in providers so a future provider can be added with just its catalog TOML. ## Changes - Replaces the closed-enum round-trip guardrail with a catalog metadata guardrail, allowing built-in TOML providers that do not have `Provider` enum variants. - Makes the all-model `fabro model` CLI tests assert stable table structure instead of snapshotting every built-in catalog row. - Renames synthetic custom-provider and missing-provider fixtures away from provider names that can become real catalog entries. ## Verification - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo nextest run -p fabro-model -p fabro-auth -p fabro-llm -p fabro-server -p fabro-workflow -p fabro-config` - `cargo nextest run -p fabro-cli cmd::model` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 (unknown context, medium reasoning) via [Codex](https://openai.com/codex) --------- Co-authored-by: Jesse <606+jesseproudman@users.noreply.github.com> |
||
|
|
e7e4fb5ca1
|
Add Daytona volume mount passthrough (#263)
## Summary This adds run-configuration support for mounting existing Daytona volumes into Fabro-managed Daytona sandboxes. Concretely, this PR: - adds `[[run.sandbox.daytona.volumes]]` with `volume_id`, `mount_path`, and optional `subpath` - resolves that config through the layer/settings/runtime pipeline - forwards configured mounts to `daytona_sdk::SandboxBaseParams.volumes` when creating the sandbox - documents the configuration surface in the Daytona environment and run configuration docs ## Motivation Daytona already supports attaching volumes when a sandbox is created, but Fabro currently owns that sandbox creation call. That means users cannot attach a pre-created Daytona volume to a Fabro-managed sandbox from run config. The intended use is persistent, provider-owned state such as agent credentials, caches, datasets, or other files that should survive ephemeral sandbox lifecycles. ## Scope This is intentionally a narrow passthrough. Fabro does not create, delete, list, wait on, or otherwise manage Daytona volume lifecycle. Users create the volume in Daytona first, then reference its `volume_id` from Fabro run config. `volumes` defaults to an empty list in resolved settings for backwards compatibility with existing serialized settings. ## Testing - `cargo test -p fabro-server runtime_daytona_config_preserves_volume_mounts` - `cargo test -p fabro-config resolves_daytona_volume_mounts` - `cargo test -p fabro-sandbox --features daytona volume_mounts` - `cargo test -p fabro-workflow runtime_daytona_config_preserves_volume_mounts` - `cargo check -p fabro-server` --- _Re-opened from #262 (originally by @kimprobably) to land a rustfmt fix — the original PR came from an org-owned fork, which blocks maintainer pushes. Branch is now on the base repo. Original commit preserved; one additional commit fixes rustfmt formatting._ Co-authored-by: Tim Keen <tim@keen.digital> |
||
|
|
c0fe29390a
|
feat(sandbox): prepare clone layout for multi-repo runs (#250)
## Summary
- Clone primary GitHub repos into provider-owned `/repos/{owner}/{repo}`
paths for Docker and Daytona sandboxes.
- Keep user/agent execution rooted at the workspace symlink, e.g.
`/workspace/{repo}` or `/home/daytona/workspace/{repo}`.
- Persist optional runtime layout metadata (`workspace_root`,
`repos_root`, `primary_repo_path`, `primary_repo_link`) through events,
projections, OpenAPI, Rust API tests, and the TS client.
- Preserve empty workspace behavior and reconnect from stored
`working_directory` for existing run records.
## Verification
- `cargo nextest run -p fabro-sandbox --features docker,daytona`
- `cargo nextest run -p fabro-workflow`
- `cargo nextest run -p fabro-server`
- `cargo build -p fabro-api`
- `cargo nextest run -p fabro-api run_sandbox_json_matches_openapi_shape
sandbox_details_json_matches_openapi_shape`
- `cd lib/packages/fabro-api-client && bun run typecheck`
- `cargo +nightly-2026-04-14 fmt --check --all`
- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D
warnings`
- `git diff --check`
## Notes
- Added ignored live smoke tests for Docker and Daytona layout
validation; they require real provider credentials/runtime.
|
||
|
|
1cfe9419f5
|
refactor(llm): simplify catalog cleanup paths (#261)
## Summary This is a cleanup pass over the configurable LLM provider/catalog work from issue #210. It addresses reuse, quality, and efficiency findings from the phase 0-9 review without changing the public provider settings contract. Notable changes: - skip LLM client initialization during run preflight when the workflow has no LLM nodes - make preflight provider checks use alias-aware `Client::has_provider` - resolve `run.model.fallbacks` through the catalog instead of the old empty-key bridge - paginate `/models` before cloning returned rows - share label parsing, provider default-adapter lookup, enum expected-value formatting, and billing token formatting helpers - use catalog provider display names for OpenAI-compatible agent profiles - align process-env configured-provider discovery with `EnvCredentialSource` ## Verification - `cargo check -p fabro-config -p fabro-model -p fabro-auth -p fabro-agent -p fabro-workflow -p fabro-server -p fabro-cli` - `cargo nextest run -p fabro-config parse_labels_keeps_key_value_pairs` - `cargo nextest run -p fabro-workflow resolve_fallback_chain_resolves` - `cargo nextest run -p fabro-auth configured_providers` - `cargo nextest run -p fabro-server list_models` - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy -p fabro-model -p fabro-auth -p fabro-workflow -p fabro-server -p fabro-agent -p fabro-config -p fabro-cli --all-targets -- -D warnings` - `cd apps/fabro-web && bun test app/routes/run-billing.test.tsx` - `cd apps/fabro-web && bun run typecheck` - `git diff --check` --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
1b6189ee32
|
docs(llm): finish configurable provider cleanup (#260)
Some checks are pending
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
TypeScript / Build (push) Waiting to run
## Summary Finish phase 9 of the configurable LLM provider/model work by aligning public docs, release notes, and guardrails with the implementation already landed in phases 0-8. - documents settings-driven providers/models, OpenAI-compatible gateway examples, typed `extra_headers`, model `api_id`, controls, and per-speed costs - adds the 2026-05-13 changelog entry and provider string migration note - updates the internal phase plan ledger to reflect current implementation status - adds a workspace policy test blocking direct production `Catalog::builtin()` usage outside catalog owner/test code - clarifies `Provider` as a built-in compatibility enum while open-ended identity is `ProviderId` ## Verification - `cargo nextest run -p fabro-dev --features dev --test it policy` - `cargo dev docs check` - `cargo nextest run -p fabro-model -p fabro-config -p fabro-auth -p fabro-llm` - `cargo build --workspace` - `cargo nextest run --workspace` (5717 passed, 182 skipped, nextest reported 1 leaky test) - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` - `git diff --check` |
||
|
|
a81eb09e78
|
feat(llm): add catalog controls and speed billing (#249)
## Summary This PR advances the catalog-driven LLM work from fabro-sh/fabro#210 by making the resolved model catalog the source of truth for provider registration, request control validation, and billing identity. Runs now preserve canonical provider/model/speed identity through pricing and API responses instead of collapsing billing around provider API aliases or model IDs alone. ## What Changed - Register LLM provider adapters from the resolved catalog, including custom OpenAI-compatible providers and their credential resolution paths. - Validate effective model request controls, including run-level defaults and node overrides, before dispatching LLM requests. - Add catalog-aware billing lookup that prices canonical `ModelRef` values, uses base model costs for standard speed, applies per-speed cost overrides, and returns an unknown estimate instead of silently billing zero for unsupported combinations. - Move Anthropic Opus fast-mode pricing into the built-in catalog for `claude-opus-4-6` and `claude-opus-4-7`. - Thread the injected catalog and effective speed controls through workflow billing, including API-mode and CLI-mode handlers. - Update billing APIs, server aggregation, generated clients, and the web billing view to expose provider/model/speed billing identity and keep standard and fast usage in separate rows. ## Notes for Review Billing lookup intentionally uses canonical catalog model IDs. Provider `api_id` substitution remains limited to provider request construction, so aliases can be used on the wire without changing billing identity. Event conversion paths that do not have catalog access now preserve token counts with a null dollar estimate rather than falling back to the bootstrap catalog. ## Verification - `cargo build -p fabro-api` - `cd lib/packages/fabro-api-client && bun run generate` - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` - `ulimit -n 4096 && cargo nextest run -p fabro-model -p fabro-workflow -p fabro-server -p fabro-api -p fabro-cli --no-fail-fast` - `ulimit -n 4096 && cargo nextest run --workspace --no-fail-fast` - `cd apps/fabro-web && bun run typecheck` - `cd apps/fabro-web && bun test` - `git diff --check` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
6297b200f7
|
refactor(run): add rich failure contract (#256)
## Summary Terminal run failures now use a first-class `RunFailure` contract so downstream consumers receive structured diagnostics instead of flat `error` / `causes` / `reason` fields. The wire shape keeps concise public messages, source-chain causes, classification, optional actor/signature data, and redacted exec output tail in one nested value. Refs fabro-sh/fabro#198 ## What Changed - Added `fabro_types::RunFailure` and changed `run.failed` to emit `properties.failure` with `final_git_commit_sha` for failed-run commit state. - Replaced `Conclusion.failure_reason` with `Conclusion.failure` while leaving stage-level `StageCompletion.failure_reason` untouched. - Updated workflow internals to preserve owned error source chains until terminal event projection, then convert them into `RunFailure.causes`. - Updated store, server, CLI, OpenAPI, and generated TypeScript client consumers to use the nested failure object. - Added serialization, OpenAPI replacement, projection, and lifecycle coverage for the new contract. ## Validation - `cargo nextest run -p fabro-api -p fabro-types -p fabro-workflow -p fabro-store -p fabro-server -p fabro-cli` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` - `cargo +nightly-2026-04-14 fmt --check --all` - `cd apps/fabro-web && bun run typecheck` - `cd apps/fabro-web && bun test` - `git diff --check` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) |
||
|
|
34d83db801
|
feat(model): support open provider catalog data (#245)
## Summary This PR moves Fabro’s provider/model catalog toward settings-driven provider identity by replacing the closed provider schema at the API/auth/model boundary with `ProviderId`, then loading built-in provider and model metadata from embedded per-provider TOML files. The immediate result is that built-ins now use the same settings-shaped catalog data that custom providers will use later, while request-serving paths still keep the existing bootstrap/default catalog behavior until the resolved-catalog plumbing lands. ## Changes - Replaces API-facing provider enum usage with string-backed `ProviderId`, including OpenAPI/progenitor replacements and regenerated TypeScript client models. - Routes model, auth, billing, CLI, server, and workflow call sites through provider IDs where they cross product identity boundaries. - Builds `Catalog` from settings-shaped provider/model data with validation for adapter keys, OpenAI-compatible `base_url`, duplicate aliases, provider defaults, disabled entries, model controls, and per-speed cost rows. - Replaces `catalog.json` with embedded provider TOML files under `lib/crates/fabro-model/src/catalog/providers/`. - Adds an explicit `fabro_model::bootstrap_catalog` hatch for setup/install paths and extends the dev policy test to keep bootstrap access contained. - Preserves public training and knowledge-cutoff labels in LLM model settings while still accepting bare TOML dates. ## Verification - `cargo nextest run -p fabro-model -p fabro-config -p fabro-api` — 416 passed - `cargo nextest run -p fabro-dev --features dev bootstrap_catalog_references_stay_in_allowlist` — 1 passed - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` - `cargo build --workspace` - `git diff --check` --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) |
||
|
|
087c9233f3
|
fix(validate): pick up sibling workflow.toml inputs for bare .fabro path (#242)
## Summary
- `fabro validate path/to/workflow.fabro` now auto-discovers a sibling
`workflow.toml` and loads its `[run.inputs]`, so templated graphs
validate the same way they do when invoked by name or by toml path.
- The discovery is opt-in to the user's specific graph: we only pick up
the sibling toml if its `[workflow].graph` resolves back to the `.fabro`
the user passed. Unrelated tomls in the same directory are ignored.
## Why
`fabro validate` is the natural fast-feedback tool for CI/pre-commit
hooks that iterate on changed `.fabro` files. Previously, a graph using
`{{ inputs.* }}` would fail with a generic MiniJinja "undefined value"
error when validated by path, even when a sibling `workflow.toml`
defined those inputs. The other two invocation forms (by name, by toml)
worked, which made the path form a usability cliff.
Fixes #195.
## Test plan
- [x] New integration test:
`bare_fabro_picks_up_sibling_workflow_toml_inputs` validates
`test/templated_inputs/workflow.fabro` (uses `{{ inputs.app_dir }}`) and
expects `Validation: OK`.
- [x] New unit tests in `fabro-config::project`:
- `resolve_workflow_path_picks_up_sibling_workflow_toml` — happy path.
- `resolve_workflow_path_ignores_sibling_toml_pointing_elsewhere` —
guard: don't apply an unrelated sibling toml.
- [x] `cargo nextest run --workspace` — 5585 tests pass.
- [x] `cargo +nightly-2026-04-14 fmt --check --all`, `clippy --workspace
--all-targets -- -D warnings` clean.
- [x] Manual: `fabro validate /tmp/fabro-issue-195/workflow.fabro`
(templated graph + sibling toml with `[run.inputs]`) prints `Validation:
OK`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Nate Aune <118984+natea@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
a33d17c88d
|
feat(run): add managed branch controls (#243)
## Summary Adds run-level controls for clone behavior, managed run branch setup/pushes, and metadata branch writes/pushes so workflows can opt out of Fabro-managed Git behavior without relying on provider-specific `skip_clone` settings. This closes fabro-sh/fabro#240. ## What Changed - Introduced `[run.clone]`, `[run.run_branch]`, and `[run.meta_branch]` settings with defaults that preserve current behavior. - Removed user-facing `skip_clone` from Docker/Daytona config while mapping the new run-level clone setting into the internal sandbox runtime options. - Gated run branch setup/push, metadata branch writer creation/push, and PR branch output on the new settings. - Enforced invalid combinations: pull requests require an enabled pushed run branch, and disabling the run branch also disables metadata branch behavior. - Updated OpenAPI, the generated TypeScript API client, frontend fixture data, and docs for the new configuration shape. ## Testing - `cargo nextest run -p fabro-config -p fabro-types -p fabro-workflow -p fabro-server` - `cargo build -p fabro-api` - `cd lib/packages/fabro-api-client && bun run generate` - `cd lib/packages/fabro-api-client && bun run typecheck` - `cd apps/fabro-web && bun run typecheck` - `cd apps/fabro-web && bun test` - `cargo build --workspace` - `cargo +nightly-2026-04-14 fmt --check --all` - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` - `cargo insta pending-snapshots` - `git diff --check` ## Post-Deploy Monitoring & Validation - Validation window: first 24 hours after release; owner: release owner/on-call engineer. - Log queries/search terms: `run_branch`, `meta_branch`, `clone.enabled`, `skip_clone`, `pull request requires an enabled pushed run branch`, `metadata branch`. - Healthy signals: runs without custom branch config continue creating and pushing run/meta branches; runs with `[run.clone] enabled = false` start provider sandboxes without cloning; runs with branch pushes disabled complete without Git push errors. - Failure signals: increased run startup failures for Docker/Daytona, unexpected PR creation conflicts, missing metadata for default-config runs, or validation errors for configurations that previously used default settings. - Mitigation trigger: if default-config runs stop producing expected branch/metadata artifacts or sandbox startup failures increase, roll back the release or temporarily restore previous defaults while investigating the run-level setting resolution path. --- [](https://github.com/EveryInc/compound-engineering-plugin) 🤖 Generated with GPT-5 via [Codex](https://openai.com/codex) Co-authored-by: Haroldo Olivieri <6575718+haroldolivieri@users.noreply.github.com> |
||
|
|
234bd5663e
|
Add ACP backend support (#237)
## Summary Implemented ACP support as a first-class Fabro backend alongside `api` and `cli`. This adds a new `fabro-acp` crate using the official ACP Rust crates, routes `backend=\"acp\"` for agent and prompt nodes, adds sandbox stdio support for local/Docker/test-support paths, emits ACP workflow events/projections, updates server steerability handling, validation, documentation, and black-box CLI coverage. ## Test Plan Passed strict non-live verification: - `ulimit -n 4096 && cargo nextest run -p fabro-workflow --run-ignored all --no-fail-fast` — 1162 passed, 0 skipped. - `ulimit -n 4096 && cargo nextest run -p fabro-acp -p fabro-sandbox -p fabro-workflow -p fabro-validate -p fabro-store -p fabro-server -p fabro-cli --run-ignored all --no-fail-fast -E 'not test(daytona_streaming_live_smoke)'` — 3125 passed. - `cargo build --workspace` — passed. - `ulimit -n 4096 && cargo nextest run --workspace --run-ignored all --no-fail-fast -E 'not test(daytona_streaming_live_smoke)'` — 5666 passed. - `cargo +nightly-2026-04-14 fmt --check --all` — passed. - `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` — passed. Live-environment tests skipped/excluded under explicit user override: - `daytona_streaming_live_smoke` was excluded from final nextest runs because it requires live Daytona infrastructure and `DAYTONA_API_KEY`. - Confirmed with `env -u DAYTONA_API_KEY cargo test -p fabro-sandbox --features daytona --test daytona_streaming_live daytona_streaming_live::daytona_streaming_live_smoke -- --ignored --exact --nocapture`: failed fast with `DAYTONA_API_KEY must be set to run this live smoke test`. |
||
|
|
a19f6dd03a
|
feat(cli): add Fabro MCP server (#236)
## Summary Adds a stdio-based Fabro MCP server so MCP clients can manage Fabro workflow runs through the authenticated `fabro` CLI, without a separate MCP auth flow. ## What Changed - Adds `fabro mcp start`, `fabro mcp config`, and `fabro mcp init <agent>` for launching and configuring the MCP server. - Introduces a new `fabro-mcp-server` crate with run-management tools: - `fabro_run_create` - `fabro_run_search` - `fabro_run_interact` - `fabro_run_gather` - `fabro_run_events` - Reuses the CLI's authenticated server connection behavior, including OAuth refresh, dev-token/local-server handling, explicit server targets, proxy behavior, and stdio env/cwd isolation. - Moves shared run-manifest construction into `fabro-manifest` so CLI runs and MCP-created runs use the same override semantics. - Extends MCP client stdio support with configured cwd and exact environment handling for reliable spawned-server tests. --------- Co-authored-by: fabro-sh-0530[bot] <281434857+fabro-sh-0530[bot]@users.noreply.github.com> Co-authored-by: Fabro <noreply@fabro.sh> |
||
|
|
b5101bbde3
|
refactor(types): remove legacy run summary shape
Use the canonical nested Run DTO directly and reject the old flat run summary JSON shape. Update store, server, CLI, and fixtures to read and produce canonical fields. |