From 0ec2d955062b867002bcf439b3df9dca7096f04d Mon Sep 17 00:00:00 2001 From: Yuneng Jiang Date: Wed, 26 Aug 2026 23:38:43 -0700 Subject: [PATCH] fix(e2e): disable thinking on the gemini chat cost test instead of racing its budget `test_gemini_chat_returns_content_and_logs_cost` asks gemini-2.5-flash to "reply with the single word pong" under `max_tokens=32`, and has been seen returning no content at all: completion_tokens=29, reasoning_tokens=29, content=None gemini-2.5-flash defaults to dynamic thinking, and `max_tokens` maps to `maxOutputTokens`, which on the 2.5 family counts thinking tokens as well as visible output. So the model is free to spend the entire budget on thoughts and emit nothing, which is exactly what the usage above shows. Raising the limit alone does not fix this. Dynamic thinking on 2.5 Flash is documented up to 24576 tokens, so no budget small enough to be reasonable for a one-word smoke test is safe. The fix is to take thinking out of the picture: `reasoning_effort="none"` maps to `thinkingConfig.thinkingBudget=0` for the 2.5 family, so the whole limit is available to visible output. Verified against this checkout: get_optional_params(model="gemini-2.5-flash", custom_llm_provider="gemini", max_tokens=32) -> {'max_output_tokens': 32} # no thinkingConfig at all get_optional_params(model="gemini-2.5-flash", custom_llm_provider="gemini", max_tokens=64, reasoning_effort="none") -> {'max_output_tokens': 64, 'thinkingConfig': {'thinkingBudget': 0, 'includeThoughts': False}} This mirrors what the OpenAI tool tests in this same file already do with gpt-5.6 for the same failure mode. `max_tokens` goes to 64 for headroom; with thinking disabled that is ample for a one-word answer. Neither `covers` claim changes: the call still exercises the gemini chat translation path and still produces a costed SpendLogs row. --- .../llm_translation/test_chat_completions_regression_e2e.py | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/tests/e2e/llm_translation/test_chat_completions_regression_e2e.py b/tests/e2e/llm_translation/test_chat_completions_regression_e2e.py index 655d426c28d..156f3393530 100644 --- a/tests/e2e/llm_translation/test_chat_completions_regression_e2e.py +++ b/tests/e2e/llm_translation/test_chat_completions_regression_e2e.py @@ -308,7 +308,8 @@ class TestGeminiChatCompletions: content=f"Reply with the single word pong. marker={tag}", ) ], - max_tokens=32, + max_tokens=64, + reasoning_effort="none", ), ) )