fabro/lib
Bryan Helmkamp 6599f7cdc0
Cover the pebble agent loop with workflow-level tests
Each behaviour the pebble backend owes a run now has one test that drives
it through the workflow engine against a scripted OpenAI-compatible model:

- an agent stage under every harness profile (openai, anthropic, claude-5,
  gemini, kimi) writing a file with that profile's own tool spelling, with
  the event sequence, files touched, response, usage, and cost checked; the
  codex vocabulary applies a patch through the OpenAI twin's custom tool
  call, which `TwinToolCall::custom` now scripts
- steering delivered mid-stage, an interrupt with a steer, run cancellation,
  and the executor-enforced stage timeout
- a question answered through the interviewer, a subagent whose events carry
  its parent's session id, and an MCP tool served by a stdio server
- failover to a second provider after a tool ran, continuing the recorded
  conversation without running the tool again
- a failing event sink ending the stage with the sink's error
- Ask Fabro resuming a stored record across turns, with the cursor moved
  past the run's event log when the record's own cursor fell behind
- Docker and Daytona smokes running an agent stage through the provider
  sandboxes

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 16:32:52 -06:00
..
apps Cover the pebble agent loop with workflow-level tests 2026-09-11 16:32:52 -06:00
components Cover the pebble agent loop with workflow-level tests 2026-09-11 16:32:52 -06:00
foundation Cover the pebble agent loop with workflow-level tests 2026-09-11 16:32:52 -06:00
packages/fabro-api-client Run agent stages, Ask Fabro, and fabro exec on pebble's CodingAgent 2026-09-11 14:19:15 -06:00