litellm/tests/unit/test_groq_streaming_encoding.py
yuneng-jiang f6882246d4
test: move tests/test_litellm root and small trees into tests/unit (#43186)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 11:30:43 -07:00

153 lines
4.8 KiB
Python

"""
Test for Groq streaming ASCII encoding issue fix.
This test verifies that the OpenAI-like handler correctly handles
UTF-8 encoded content in streaming responses, specifically fixing
the ASCII encoding error described in issue #12660.
"""
import asyncio
from unittest.mock import AsyncMock, Mock
import pytest
from litellm.llms.openai_like.chat.handler import make_call, make_sync_call
class MockResponse:
"""Mock httpx response for testing UTF-8 handling."""
def __init__(self, test_content: str):
self.test_content = test_content
self.status_code = 200
def iter_text(self, encoding="utf-8"):
"""Mock iter_text that yields content with the specified encoding."""
yield self.test_content
async def aiter_text(self, encoding="utf-8"):
"""Mock aiter_text that yields content with the specified encoding."""
yield self.test_content
def iter_lines(self):
"""Mock iter_lines method for synchronous streaming."""
yield self.test_content
async def aiter_lines(self):
"""Mock aiter_lines method for asynchronous streaming."""
yield self.test_content
def json(self):
return {"choices": [{"delta": {"content": "test"}}]}
class MockSyncClient:
"""Mock synchronous HTTP client for testing."""
def __init__(self, response_content: str):
self.response_content = response_content
def post(self, *args, **kwargs):
return MockResponse(self.response_content)
class MockAsyncClient:
"""Mock asynchronous HTTP client for testing."""
def __init__(self, response_content: str):
self.response_content = response_content
async def post(self, *args, **kwargs):
return MockResponse(self.response_content)
def test_utf8_streaming_sync():
"""Test that synchronous streaming handles UTF-8 characters correctly."""
# Content with the µ character that was causing issues
test_content = (
'data: {"choices":[{"delta":{"content":"The symbol µ represents micro"}}]}\n\n'
)
mock_client = MockSyncClient(test_content)
mock_logging = Mock()
# This should not raise an ASCII encoding error
completion_stream = make_sync_call(
client=mock_client,
api_base="https://test.com/v1/chat/completions",
headers={"Authorization": "Bearer test"},
data='{"model": "test", "messages": []}',
model="test-model",
messages=[],
logging_obj=mock_logging,
)
# Verify we can iterate through the stream without encoding errors
assert completion_stream is not None
@pytest.mark.asyncio
async def test_utf8_streaming_async():
"""Test that asynchronous streaming handles UTF-8 characters correctly."""
# Content with the µ character that was causing issues
test_content = (
'data: {"choices":[{"delta":{"content":"The symbol µ represents micro"}}]}\n\n'
)
mock_client = MockAsyncClient(test_content)
mock_logging = Mock()
# This should not raise an ASCII encoding error
completion_stream = await make_call(
client=mock_client,
api_base="https://test.com/v1/chat/completions",
headers={"Authorization": "Bearer test"},
data='{"model": "test", "messages": []}',
model="test-model",
messages=[],
logging_obj=mock_logging,
)
# Verify we can iterate through the stream without encoding errors
assert completion_stream is not None
def test_various_unicode_characters():
"""Test streaming with various Unicode characters that could cause issues."""
unicode_test_cases = [
"µ", # Micro symbol (the original issue)
"©", # Copyright symbol
"™", # Trademark symbol
"€", # Euro symbol
"北京", # Chinese characters
"🚀", # Emoji
"Ñoño", # Spanish characters with tildes
]
for unicode_char in unicode_test_cases:
test_content = f'data: {{"choices":[{{"delta":{{"content":"Testing {unicode_char} character"}}}}]}}\n\n'
mock_client = MockSyncClient(test_content)
mock_logging = Mock()
# This should not raise an ASCII encoding error for any Unicode character
completion_stream = make_sync_call(
client=mock_client,
api_base="https://test.com/v1/chat/completions",
headers={"Authorization": "Bearer test"},
data='{"model": "test", "messages": []}',
model="test-model",
messages=[],
logging_obj=mock_logging,
)
assert (
completion_stream is not None
), f"Failed to handle Unicode character: {unicode_char}"
if __name__ == "__main__":
test_utf8_streaming_sync()
asyncio.run(test_utf8_streaming_async())
test_various_unicode_characters()
print("All UTF-8 streaming tests passed!")