litellm/tests/test_litellm/llms/openai
Cesar Garcia 75d0d2bd7a fix(openrouter): preserve token counts from streaming usage chunks (#21011)
* docs: add reference to example_openai_endpoint repo for self-hosting fake OpenAI proxy (#21006)

- Updated benchmarks.md with a section on setting up fake OpenAI endpoints
- Updated load_test.md to mention the self-hosted option
- Updated load_test_advanced.md with a tip box about the example repo

Reference: https://github.com/BerriAI/example_openai_endpoint

Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* MCP fixes

* fix(oldteams.tsx): show policies when creating

* fix(proxy/_types.py): ensure mcp rest endpoints can be called by virtual key

ensures UI works with virtual key testing mcp endpoints

* refactor: migrate get object permissions table logic to happen in user api key auth - allows functions to trust user api key object they receive has what they need

* fix(rest_endpoints.py): filter for allowed tools based on what key has access to

* fix(mcp_server_manager.py): ensure only allowed MCP's are returned to the user, via rest endpoints

* Guardrails - add toxic/abusive content filter guardrails

* fix(streaming): preserve usage data from post-finish_reason chunks in OpenAI-compatible streaming

Fixes #16112

OpenRouter and other OpenAI-compatible providers send a usage chunk after
the finish_reason='stop' chunk when stream_options.include_usage is True.
The OpenAIChatCompletionStreamingHandler.chunk_parser() was not passing
the usage field to ModelResponseStream, causing real token counts from the
provider to be lost and falling back to inaccurate estimates.

* fix: resolve merge conflict in test file

- Fix typo in test method name (extra space)
- Move test_prompt_cache_key_in_optional_params to its own class

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-02-13 18:27:22 +05:30
..
chat fix(openrouter): preserve token counts from streaming usage chunks (#21011) 2026-02-13 18:27:22 +05:30
completion fix(text_completion): support token IDs (list of integers) as prompt (#18011) 2026-01-12 17:33:24 +05:30
embeddings/guardrail_translation [Fix] Guardrails API - Ensure OpenAI Moderations Guard works with OpenAI Embeddings (#20523) 2026-02-05 14:40:15 -08:00
image_generation fix(unified_guardrail.py): support during_call event type for unified guardrails (#17514) 2025-12-04 22:06:13 -08:00
realtime [Feat] Add xAI /realtime API Support - works with LiveKitSDK (#20381) 2026-02-03 19:58:28 -08:00
responses fix(responses): handle Pydantic ValidationError when provider omits required fields in streaming events (#20580) 2026-02-12 19:55:58 +05:30
speech fix(unified_guardrail.py): support during_call event type for unified guardrails (#17514) 2025-12-04 22:06:13 -08:00
transcriptions fix(unified_guardrail.py): support during_call event type for unified guardrails (#17514) 2025-12-04 22:06:13 -08:00
vector_store_files Vector store files Stable Release (#16643) 2025-11-15 13:00:33 -08:00
vector_stores Fix create, search vector store error (#13285) 2025-08-06 11:15:17 -07:00
test_gpt5_transformation.py fix: allow tool_choice for Azure GPT-5 chat models (#19813) 2026-01-27 17:51:13 -08:00
test_o_series_transformation.py 🐛 Fix a bug where openai image edit siltently ignores multiple images 2025-09-26 10:56:06 +05:30
test_openai_common_utils.py Litellm fix GitHub action testing (#11163) 2025-05-26 14:41:42 -07:00
test_openai_empty_response.py (fix): empty response + vllm streaming (#17516) 2025-12-04 19:24:28 -08:00
test_openai_image_edit_transformation.py fix(openai/image_edit/transformation.py): fix passing multiple images to openai 2025-09-27 16:03:17 -07:00