litellm/tests
Ishaan Jaff 9761ba7c7a
[Bug Fix] Responses api session management for streaming responses (#13396)
* fix proxy config

* fix(responses api): fix streaming ID consistency and tool format handling (#12640)

* fix(responses): ensure streaming chunk IDs use consistent encoding format

Fixes streaming ID inconsistency where streaming responses used raw provider IDs
while non-streaming responses used properly encoded IDs with provider context.

Changes:
- Updated LiteLLMCompletionStreamingIterator to accept provider context
- Added _encode_chunk_id() method using same logic as non-streaming responses
- Modified chunk transformation to encode all streaming item_ids with resp_ prefix
- Updated handlers to pass custom_llm_provider and litellm_metadata to streaming iterator

Impact:
- Streaming chunk IDs now format: resp_<base64_encoded_provider_context>
- Enables session continuity when using streaming response IDs as previous_response_id
- Allows provider detection and load balancing with streaming responses
- Maintains backward compatibility with existing streaming functionality

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(types): add explicit Optional[str] type annotation for model_id

This resolves MyPy type checking error where model_id could be None
but wasn't explicitly typed as Optional[str].

* fix(types): handle None case for litellm_metadata access

Prevents 'Item None has no attribute get' error by checking for None
before accessing litellm_metadata dictionary.

* test: add comprehensive tests for streaming ID consistency

Adds unit and E2E tests to verify streaming chunk IDs are properly encoded
with consistent format across streaming responses.

## Tests Added

### Unit Test (test_reasoning_content_transformation.py)
- `test_streaming_chunk_id_encoding()`: Validates the `_encode_chunk_id()` method
  correctly encodes chunk IDs with `resp_` prefix and provider context

### E2E Tests (test_e2e_openai_responses_api.py)
- `test_streaming_id_consistency_across_chunks()`: Tests that all streaming chunk IDs
  are properly encoded across multiple chunks in a real streaming response
- `test_streaming_response_id_as_previous_response_id()`: Tests the core use case -
  using streaming response IDs for session continuity with `previous_response_id`

## Key Testing Approach
- Uses **Gemini** (non-OpenAI model) to test the transformation logic rather than
  OpenAI passthrough, since the streaming ID consistency issue occurs when LiteLLM
  transforms responses rather than just passing through to native OpenAI responses API
- Tests validate that streaming chunk IDs now use same encoding as non-streaming responses
- Verifies session continuity works with streaming responses

Addresses @ishaan-jaff's request for unit tests covering the streaming ID consistency fix.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(lint): remove unused imports in transformation.py

Removes unused imports to fix CI linting errors:
- GenericResponseOutputItem
- OutputFunctionToolCall

* test: remove E2E tests from openai_endpoints_tests

Remove streaming ID consistency E2E tests as requested by @ishaan-jaff.
Keep only the mock/unit test in test_reasoning_content_transformation.py

* revert: remove streaming chunk ID encoding to original behavior

This reverts the streaming chunk ID encoding changes to understand the original issue better.
Original behavior was:
- Streaming chunks: raw provider IDs
- Streaming final response: raw IDs (PROBLEM!)
- Non-streaming final response: encoded IDs (correct)

The real issue: streaming final response IDs were not encoded, breaking session continuity.

* fix(responses): encode streaming final response IDs to match OpenAI behavior

Fixes streaming ID inconsistency to match OpenAI's Responses API behavior:
- Streaming chunks: raw message IDs (like OpenAI's msg_xxx)
- Final response: encoded IDs (like OpenAI's resp_xxx)

This enables session continuity by ensuring streaming final response IDs
have the same encoded format as non-streaming responses, allowing them
to be used as previous_response_id in follow-up requests.

Changes:
- Add custom_llm_provider and litellm_metadata to LiteLLMCompletionStreamingIterator
- Update handlers to pass provider context to streaming iterator
- Apply _update_responses_api_response_id_with_model_id to final streaming response
- Keep streaming chunks as raw IDs to match OpenAI format

Impact:
- Session continuity works with streaming responses
- Load balancing can detect provider from streaming final response IDs
- Format matches OpenAI's Responses API exactly

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>

* test: update unit test to match correct OpenAI-compatible behavior

Updates the unit test to verify streaming chunk IDs are raw (not encoded)
to match OpenAI's responses API format:
- Streaming chunks: raw message IDs (like msg_xxx)
- Final response: encoded IDs (like resp_xxx)

This reflects the correct behavior implemented in the fix.

---------

Co-authored-by: Claude <noreply@anthropic.com>

* cleanup

* TestBaseResponsesAPIStreamingIterator

---------

Co-authored-by: Javier de la Torre <jatorre@carto.com>
Co-authored-by: Claude <noreply@anthropic.com>
2025-08-07 20:13:24 -07:00
..
basic_proxy_startup_tests (fix) don't block proxy startup if license check fails & using prometheus (#6839) 2024-11-20 17:55:39 -08:00
batches_tests openai.ConflictError 2025-07-12 17:07:21 -07:00
code_coverage_tests [Bug Fix] Mistral Tool Calling - Grammar error: at 3(11): failed to compile JSON schema (#13389) 2025-08-07 13:50:22 -07:00
documentation_tests Litellm dev 12 28 2024 p1 (#7463) 2024-12-28 20:26:00 -08:00
enterprise [Feat] Responses API Session Handling - Multi media support (#13347) 2025-08-07 10:59:53 -07:00
guardrails_tests feat: Add Pillar Security guardrail integration (#12791) 2025-07-22 23:17:33 -07:00
image_gen_tests [Feat] Add Google AI Studio Imagen4 model family (#13065) 2025-07-28 21:25:40 -07:00
litellm-proxy-extras Prisma Migrate - support setting custom migration dir (#10336) 2025-04-26 12:05:06 -07:00
litellm_utils_tests test: update test 2025-07-29 21:08:36 -07:00
llm_responses_api_testing [Bug Fix] Responses api session management for streaming responses (#13396) 2025-08-07 20:13:24 -07:00
llm_translation [Bug Fix] Mistral Tool Calling - Grammar error: at 3(11): failed to compile JSON schema (#13389) 2025-08-07 13:50:22 -07:00
load_tests test: test_embedding_performance 2025-05-14 21:31:07 -07:00
local_testing test_stream_chunk_builder_litellm_usage_chunks 2025-08-07 15:22:52 -07:00
logging_callback_tests [Feat] Responses API Session Handling - Multi media support (#13347) 2025-08-07 10:59:53 -07:00
mcp_tests [MCP Gateway] fix auth on ui for bearer servers (#13312) 2025-08-06 09:46:10 -07:00
multi_instance_e2e_tests (e2e testing) - add tests for using litellm /team/ updates in multi-instance deployments with Redis (#8440) 2025-02-10 19:33:27 -08:00
old_proxy_tests/tests test: update tests to new deployment model (#10142) 2025-04-18 14:22:12 -07:00
openai_endpoints_tests Batches - support batch retrieve with target model Query Param + Anthropic - completion bridge, yield content_block_stop chunk (#12228) 2025-07-01 22:13:48 -07:00
otel_tests test_bedrock_guardrail_triggered 2025-07-09 17:05:06 -07:00
pass_through_tests test_anthropic_basic_completion_with_headers ci/cd fix 2025-07-12 11:08:03 -07:00
pass_through_unit_tests Azure api_version="preview" support + Bedrock cost tracking via Anthropic /v1/messages (#13072) 2025-07-29 08:13:55 -07:00
proxy_admin_ui_tests Fix e2e test (#12549) 2025-07-12 10:42:57 -07:00
proxy_security_tests (Security fix) - remove code block that inserts master key hash into DB (#8268) 2025-02-05 17:25:42 -08:00
proxy_unit_tests Exclude none fields on /chat/completion - fixes n8n bug + Allow calling /v1/models when end user over budget (#13320) 2025-08-05 21:39:46 -07:00
router_unit_tests [LLM Translation] Fix model group on clientside auth with API calls (#13314) 2025-08-05 17:46:47 -07:00
scim_tests [Feat SSO] Add LiteLLM SCIM Integration for Team and User management (#10072) 2025-04-16 19:21:47 -07:00
spend_tracking_tests test_long_term_spend_accuracy_with_bursts 2025-03-31 21:09:29 -07:00
store_model_in_db_tests [MCP Gateway] fix auth on ui for bearer servers (#13312) 2025-08-06 09:46:10 -07:00
test_litellm [Bug Fix] Responses api session management for streaming responses (#13396) 2025-08-07 20:13:24 -07:00
unified_google_tests test fix: gcp deprecated gemini-1.5-flash 2025-08-06 08:43:45 -07:00
vector_store_tests [Feat] UI Vector Stores - Allow adding Vertex RAG Engine, OpenAI, Azure (#12752) 2025-07-18 18:25:26 -07:00
windows_tests [Bug Fix] UnicodeDecodeError: 'charmap' on Windows during litellm import (#10542) 2025-05-03 21:31:05 -07:00
__init__.py [Feat] Add github co-pilot as a new LLM API provider (#12325) 2025-07-04 13:12:16 -07:00
gettysburg.wav feat(main.py): support openai transcription endpoints 2024-03-08 10:25:19 -08:00
large_text.py fix(router.py): check for context window error when handling 400 status code errors 2024-03-26 08:08:15 -07:00
openai_batch_completions.jsonl feat(router.py): Support Loadbalancing batch azure api endpoints (#5469) 2024-09-02 21:32:55 -07:00
README.MD add bedrock llama vision support + cohere / infinity rerank - 'return_documents' support (#8684) 2025-02-20 21:23:54 -08:00
test_budget_management.py Update enduser spend and budget reset date based on budget duration (#8460) 2025-06-08 08:39:14 -07:00
test_callbacks_on_proxy.py fix - test num callbacks 2024-05-17 22:06:51 -07:00
test_config.py fix testing - langfuse apis are flaky, we unit test team / key based logging in test_langfuse_unit_tests.py 2024-12-03 11:24:36 -08:00
test_debug_warning.py fix(utils.py): fix togetherai streaming cost calculation 2024-08-01 15:03:08 -07:00
test_end_users.py test: run test earlier to catch error 2025-03-27 23:08:52 -07:00
test_entrypoint.py (fix) clean up root repo - move entrypoint.sh and build_admin_ui to /docker (#6110) 2024-10-08 11:34:43 +05:30
test_fallbacks.py Ollama Chat - parse tool calls on streaming (#11171) 2025-05-27 16:14:49 -07:00
test_health.py (test) /health/readiness 2024-01-29 15:27:25 -08:00
test_keys.py test: temporarily skip test due to change testing model change - need to update test for new model 2025-05-09 09:02:08 -07:00
test_logging.conf feat(proxy_cli.py): add new 'log_config' cli param (#6352) 2024-10-21 21:25:58 -07:00
test_models.py Add model access groups on UI (#11719) 2025-06-13 21:20:25 -07:00
test_openai_endpoints.py test_chat_completion_ratelimit 2025-07-30 23:12:11 -07:00
test_organizations.py UI - fix adding vertex models with reusable credentials + fix pagination on keys table + fix showing org budgets on table (#10528) 2025-05-03 08:16:53 -07:00
test_passthrough_endpoints.py test test_basic_passthrough 2024-08-06 21:17:07 -07:00
test_ratelimit.py (Refactor / QA) - Use LoggingCallbackManager to append callbacks and ensure no duplicate callbacks are added (#8112) 2025-01-30 19:35:50 -08:00
test_resource_cleanup.py Fix: Properly close aiohttp client sessions to prevent resource leaks (#12251) 2025-07-09 09:25:17 -07:00
test_spend_logs.py (feat) - track org_id in SpendLogs (#8253) 2025-02-04 21:08:05 -08:00
test_team.py build: publish new litellm-proxy-extras file 2025-05-27 17:44:23 -07:00
test_team_logging.py test: skip flaky test 2024-11-22 19:23:36 +05:30
test_team_members.py test: add more unit testing for team member endpoints (#8170) 2025-02-01 11:23:00 -08:00
test_users.py test: fix imports 2025-05-26 22:06:53 -07:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/litellm

This folder can only run mock tests.