fabro/test/twin/openai/tests/responses_contract.rs
Scott Werner e4a85679bf
fix(test): make twin-openai streamed message items round-trip as input (#484)
## What

Fixes the `openai_twin_*` parity-matrix failures that have been on
`main` since #449: every multi-turn scenario whose scripted response
includes text fails on its second turn with 400 `"message input items
require supported content"`.

## Root cause

Two twin behaviors collided (bisected: passes at #447, fails at #449):

1. **The twin's streaming `response.output_item.done` for message items
omitted the `content` array** (`test/twin/openai/src/sse.rs`) — it sent
only `id`/`type`/`status`/`role`, where the real API sends the completed
item in full. The openai adapter preserves message output items verbatim
(`ContentPart::Other { kind: OPENAI_MESSAGE }`) and replays them as
assistant history on the next turn — required so reasoning items keep
their "required following item" in Responses round-trips. So the replay
arrived content-less.
2. **#449 tightened the twin's input validation** to also validate
explicit `type: "message"` items (previously only type-less items were
validated as messages; anything with an explicit type was accepted
unchecked). The twin started rejecting its own round-tripped output.

The new validation caught a real infidelity in the emitter — the emit
side is what's wrong.

Nobody noticed because **CI never runs the twin e2e suites**: `rust.yml`
runs `--profile ci` without `--run-ignored`, so the parity matrix only
runs when someone invokes the e2e profile locally.

## Fix

- The streamed message `output_item.done` now carries its `output_text`
content, matching the real API and the twin's own non-streaming
`responses_json()`.
- The input validator accepts `output_text` parts on **assistant**
message items (the real API allows these; the twin's non-streaming
responses already require it for faithful replay). Non-assistant
`output_text` parts get a dedicated rejection message.

## Tests

- New contract test
`responses_stream_message_item_done_round_trips_as_input`: streams a
response, asserts the completed message item carries its `output_text`
content, and replays the item verbatim as assistant-history input,
asserting the twin accepts its own output.
- `cargo nextest run -p twin-openai` — 56 passed
- `cargo nextest run -p fabro-agent -E 'test(parity)' --run-ignored
only` — **91/91 passed** (was 7 failing)
- `cargo nextest run -p fabro-llm --run-ignored only` — 10 passed
- `cargo nextest run --workspace` — green apart from two pre-existing
env-dependent `fabro-workflow` failures that reproduce on clean `main`
in shells with provider API keys exported (unrelated; CI is green on
them because it has no such keys)
- clippy `-D warnings` / fmt — clean

Found while reviewing #481 (whose parity runs surfaced this); #481
itself is unaffected — it doesn't touch the openai adapter or the twin,
and the failures exist on its merge-base.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

## CI (separate commit, drop if unwanted)

`ci: run twin-mode e2e suites on Linux` adds a step to the existing
Linux test job running the ignored twin-mode suites for the packages
that are fully green today (`fabro-agent`, `fabro-llm`, `twin-openai`) —
104 tests, ~1s on a warm build, no secrets needed (live-only tests
self-skip in twin mode). This is what would have caught the #449
regression. The remaining ignored suites (fabro-cli twin tests,
Docker/Daytona sandbox tests, fabro-spa asset test) need their own fixes
before joining; widen the `-E` filter as they're cleaned up. Note the
step deliberately avoids the `e2e` nextest profile, since
`NEXTEST_PROFILE=e2e` implies strict mode, which fails on missing
secrets.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 12:56:20 -04:00

440 lines
13 KiB
Rust

mod common;
use serde_json::json;
#[tokio::test]
async fn responses_create_returns_deterministic_non_stream_payload() {
let server = common::spawn_server().await.expect("server should start");
let response = server
.post_responses_with_headers(
json!({
"model": "gpt-test",
"input": "Hello from the test suite",
"stream": false
}),
Some("org-test"),
Some("proj-test"),
)
.await;
assert_eq!(response.status(), 200);
let body = response
.json::<serde_json::Value>()
.await
.expect("json body should parse");
assert_eq!(body["object"], "response");
assert_eq!(body["model"], "gpt-test");
assert_eq!(body["status"], "completed");
assert_eq!(body["id"], "resp_000001");
assert_eq!(body["created"], 1);
assert_eq!(body["output"][0]["type"], "message");
assert_eq!(
body["output"][0]["content"][0]["text"],
"deterministic: Hello from the test suite"
);
assert_eq!(body["usage"]["input_tokens"], 5);
assert_eq!(body["usage"]["output_tokens"], 5);
}
#[tokio::test]
async fn responses_accepts_supported_openai_request_fields() {
let server = common::spawn_server().await.expect("server should start");
let response = server
.post_responses(json!({
"model": "gpt-test",
"stream": false,
"input": [
{
"role": "user",
"content": [
{ "type": "input_text", "text": "Summarize this image please" },
{ "type": "input_image", "image_url": "https://example.com/cat.png" }
]
}
],
"metadata": {
"suite": "responses",
"case": "supported-fields"
},
"stop": ["END"],
"previous_response_id": "resp_previous",
"reasoning": {
"effort": "medium"
},
"tools": [
{
"type": "function",
"name": "lookup_weather",
"description": "Look up weather",
"parameters": {
"type": "object",
"properties": {
"city": { "type": "string" }
}
}
}
],
"tool_choice": "auto",
"text": {
"format": {
"type": "text"
}
}
}))
.await;
assert_eq!(response.status(), 200);
let body = response
.json::<serde_json::Value>()
.await
.expect("json body should parse");
assert_eq!(body["object"], "response");
assert_eq!(
body["output"][0]["content"][0]["text"],
"deterministic: Summarize this image please"
);
}
#[tokio::test]
async fn responses_reject_unfulfilled_tool_choice_requirements() {
let server = common::spawn_server().await.expect("server should start");
let required = server
.post_responses(json!({
"model": "gpt-test",
"input": "plain text please",
"stream": false,
"tools": [{ "type": "function", "name": "lookup_weather" }],
"tool_choice": "required"
}))
.await;
assert_eq!(required.status(), 400);
let body = required.json::<serde_json::Value>().await.expect("json");
assert_eq!(body["error"]["type"], "invalid_request_error");
assert_eq!(body["error"]["param"], "tool_choice");
let named = server
.post_responses(json!({
"model": "gpt-test",
"input": "plain text please",
"stream": false,
"tools": [{ "type": "function", "name": "lookup_weather" }],
"tool_choice": {
"type": "function",
"name": "lookup_weather"
}
}))
.await;
assert_eq!(named.status(), 400);
let body = named.json::<serde_json::Value>().await.expect("json");
assert_eq!(body["error"]["type"], "invalid_request_error");
assert_eq!(body["error"]["param"], "tool_choice");
}
#[tokio::test]
async fn responses_accept_unknown_top_level_fields() {
let server = common::spawn_server().await.expect("server should start");
let response = server
.post_responses(json!({
"model": "gpt-test",
"input": "hello",
"stream": false,
"unexpected_field": true
}))
.await;
assert_eq!(response.status(), 200);
}
#[tokio::test]
async fn responses_reject_malformed_input_items() {
let server = common::spawn_server().await.expect("server should start");
let missing_call_id = server
.post_responses(json!({
"model": "gpt-test",
"input": [
{
"type": "function_call_output",
"output": "72 and sunny"
}
],
"stream": false
}))
.await;
assert_eq!(missing_call_id.status(), 400);
let body = missing_call_id
.json::<serde_json::Value>()
.await
.expect("json");
assert_eq!(body["error"]["type"], "invalid_request_error");
assert_eq!(body["error"]["param"], "input");
let missing_content = server
.post_responses(json!({
"model": "gpt-test",
"input": [
{
"role": "user"
}
],
"stream": false
}))
.await;
assert_eq!(missing_content.status(), 400);
let body = missing_content
.json::<serde_json::Value>()
.await
.expect("json");
assert_eq!(body["error"]["type"], "invalid_request_error");
assert_eq!(body["error"]["param"], "input");
}
#[tokio::test]
async fn responses_reject_malformed_image_input() {
let server = common::spawn_server().await.expect("server should start");
let response = server
.post_responses(json!({
"model": "gpt-test",
"input": [{
"role": "user",
"content": [{
"type": "input_image",
"image_url": ""
}]
}],
"stream": false
}))
.await;
assert_eq!(response.status(), 400);
let body = response.json::<serde_json::Value>().await.expect("json");
assert_eq!(body["error"]["type"], "invalid_request_error");
assert_eq!(body["error"]["param"], "input");
}
#[tokio::test]
async fn responses_stream_message_item_done_round_trips_as_input() {
let server = common::spawn_server().await.expect("server should start");
let (status, chunks) = server
.post_responses_stream(json!({
"model": "gpt-test",
"input": "stream this request",
"stream": true
}))
.await;
assert_eq!(status, 200);
let joined = chunks.join("");
let transcript = common::parse_sse_transcript(joined.as_bytes()).expect("valid sse");
let message_item = transcript
.events
.iter()
.filter(|event| event.event.as_deref() == Some("response.output_item.done"))
.filter_map(|event| serde_json::from_str::<serde_json::Value>(&event.data).ok())
.map(|payload| payload["item"].clone())
.find(|item| item["type"] == "message")
.expect("stream should emit a completed message item");
// The completed item carries its full content, like the real API.
let content = message_item["content"]
.as_array()
.expect("completed message item should include content");
assert_eq!(content.len(), 1);
assert_eq!(content[0]["type"], "output_text");
assert!(
content[0]["text"]
.as_str()
.is_some_and(|text| !text.is_empty())
);
// Adapters replay the completed item verbatim as assistant history on the
// next turn, so the twin must accept its own streamed output as input.
let response = server
.post_responses(json!({
"model": "gpt-test",
"input": [
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "stream this request"}]
},
message_item,
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "and again"}]
},
],
"stream": false
}))
.await;
assert_eq!(response.status(), 200);
}
#[tokio::test]
async fn responses_stream_emits_expected_sse_sequence() {
let server = common::spawn_server().await.expect("server should start");
let request = json!({
"model": "gpt-test",
"input": "stream this request",
"stream": true
});
let non_stream = server
.post_responses(json!({
"model": "gpt-test",
"input": "stream this request",
"stream": false
}))
.await
.json::<serde_json::Value>()
.await
.expect("json body should parse");
let (status, chunks) = server.post_responses_stream(request).await;
let joined = chunks.join("");
let transcript = common::parse_sse_transcript(joined.as_bytes()).expect("valid sse");
let events = transcript
.events
.iter()
.filter_map(|event| event.event.as_deref())
.collect::<Vec<_>>();
assert_eq!(status, 200);
assert_eq!(events, vec![
"response.created",
"response.in_progress",
"response.output_item.added",
"response.output_item.done",
"response.output_item.added",
"response.content_part.added",
"response.output_text.delta",
"response.output_text.done",
"response.content_part.done",
"response.output_item.done",
"response.completed",
]);
assert!(!transcript.done);
assert!(joined.contains("deterministic: stream this request"));
assert!(
joined.contains(
non_stream["output"][0]["content"][0]["text"]
.as_str()
.expect("text")
)
);
}
#[tokio::test]
async fn responses_stream_emits_reasoning_and_completion_events() {
let server = common::spawn_server().await.expect("server should start");
let (status, chunks) = server
.post_responses_stream(json!({
"model": "gpt-test",
"input": "show your reasoning",
"stream": true,
"reasoning": {
"effort": "high"
}
}))
.await;
let joined = chunks.join("");
assert_eq!(status, 200);
assert!(joined.contains("event: response.reasoning.delta\n"));
assert!(joined.contains("reasoning: show your reasoning"));
assert!(joined.contains("event: response.completed\n"));
}
#[tokio::test]
async fn responses_stream_emits_structured_output_events() {
let server = common::spawn_server().await.expect("server should start");
let non_stream = server
.post_responses(json!({
"model": "gpt-test",
"input": "structured stream",
"stream": false,
"text": {
"format": { "type": "json_object" }
}
}))
.await
.json::<serde_json::Value>()
.await
.expect("json");
let (status, chunks) = server
.post_responses_stream(json!({
"model": "gpt-test",
"input": "structured stream",
"stream": true,
"text": {
"format": { "type": "json_object" }
}
}))
.await;
let joined = chunks.join("");
let transcript = common::parse_sse_transcript(joined.as_bytes()).expect("valid sse");
let events = transcript
.events
.iter()
.filter_map(|event| event.event.as_deref())
.collect::<Vec<_>>();
let streamed_json = transcript
.events
.iter()
.find(|event| event.event.as_deref() == Some("response.output_text.done"))
.and_then(|event| serde_json::from_str::<serde_json::Value>(&event.data).ok())
.and_then(|payload| {
payload
.get("text")
.and_then(serde_json::Value::as_str)
.map(ToOwned::to_owned)
})
.and_then(|text| serde_json::from_str::<serde_json::Value>(&text).ok())
.expect("structured stream output text");
assert_eq!(status, 200);
assert_eq!(events, vec![
"response.created",
"response.in_progress",
"response.output_item.added",
"response.output_item.done",
"response.output_item.added",
"response.content_part.added",
"response.output_text.delta",
"response.output_text.done",
"response.content_part.done",
"response.output_item.done",
"response.completed",
]);
assert!(!transcript.done);
assert_eq!(
streamed_json["message"],
non_stream["output"][0]["content"][1]["json"]["message"]
);
assert!(
joined.contains(
non_stream["output"][0]["content"][1]["json"]["message"]
.as_str()
.expect("json message")
)
);
}