litellm/tests/test_litellm/llms/fireworks_ai
mateo f9c5be8ebf fix(fireworks_ai): correct Kimi K2.5/K2.6/K2.7 max output token limits
Fireworks publishes a 262144-token context window for the Kimi K2.5, K2.6
and K2.7 models but caps generation well below that. Every fireworks_ai
Kimi K2.5/K2.6/K2.7 alias had max_output_tokens/max_tokens flattened to
262144 (equal to the context window), so the pre-call context-window check
admitted requests asking for a full 262144-token completion that Fireworks
rejects. Correct max_output_tokens/max_tokens to 32768 while keeping
max_input_tokens at 262144, and add a regression test pinning the limits
for all ten aliases.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-30 01:43:59 +00:00
..
chat fix(fireworks_ai): restore Content-Type application/json header (fixes 415) (#33929) 2026-07-20 09:52:35 -07:00
rerank style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_fireworks_ai_cost_calculator.py fix(fireworks_ai): correct glm-5p2 prompt-cache read price to $0.14/1M 2026-07-17 17:57:36 -07:00
test_fireworks_ai_kimi_model_metadata.py fix(fireworks_ai): correct Kimi K2.5/K2.6/K2.7 max output token limits 2026-07-30 01:43:59 +00:00