litellm/tests/unit/test_count_tokens_public_api.py
yuneng-jiang f6882246d4
test: move tests/test_litellm root and small trees into tests/unit (#43186)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 11:30:43 -07:00

159 lines
5.2 KiB
Python

"""
Tests for litellm.acount_tokens() public API.
"""
import asyncio
import os
from unittest.mock import AsyncMock, patch
import litellm
from litellm.types.utils import TokenCountResponse
def test_acount_tokens_routes_to_openai():
"""Test that acount_tokens routes to OpenAI token counter for openai/ models."""
with patch(
"litellm.llms.openai.responses.count_tokens.token_counter.openai_count_tokens_handler.handle_count_tokens_request",
new_callable=AsyncMock,
return_value={"input_tokens": 15},
):
result = asyncio.run(
litellm.acount_tokens(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "Hello, how are you?"}],
api_key="sk-test-key",
)
)
assert result.total_tokens == 15
assert result.tokenizer_type == "openai_api"
assert result.request_model == "openai/gpt-4o"
def test_acount_tokens_routes_to_anthropic():
"""Test that acount_tokens routes to Anthropic token counter for anthropic/ models."""
with patch(
"litellm.llms.anthropic.count_tokens.token_counter.anthropic_count_tokens_handler.handle_count_tokens_request",
new_callable=AsyncMock,
return_value={"input_tokens": 20},
):
result = asyncio.run(
litellm.acount_tokens(
model="anthropic/claude-3-5-sonnet-20241022",
messages=[{"role": "user", "content": "Hello Claude!"}],
api_key="sk-ant-test-key",
)
)
assert result.total_tokens == 20
assert result.tokenizer_type == "anthropic_api"
assert result.request_model == "anthropic/claude-3-5-sonnet-20241022"
def test_acount_tokens_fallback_to_local():
"""Test that unsupported providers fall back to local tiktoken counting."""
result = asyncio.run(
litellm.acount_tokens(
model="together_ai/meta-llama/Llama-3-8b-chat-hf",
messages=[{"role": "user", "content": "Hello"}],
)
)
assert result.total_tokens > 0
assert result.tokenizer_type == "local_tokenizer"
def test_acount_tokens_with_tools():
"""Test that tools are passed through to the token counter."""
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather info",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
},
},
}
]
with patch(
"litellm.llms.openai.responses.count_tokens.token_counter.openai_count_tokens_handler.handle_count_tokens_request",
new_callable=AsyncMock,
return_value={"input_tokens": 30},
) as mock_handler:
result = asyncio.run(
litellm.acount_tokens(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "What's the weather?"}],
tools=tools,
api_key="sk-test-key",
)
)
assert result.total_tokens == 30
mock_handler.assert_called_once()
call_kwargs = mock_handler.call_args
assert call_kwargs.kwargs.get("tools") == tools
def test_acount_tokens_with_system():
"""Test that system messages are passed through."""
with patch(
"litellm.llms.openai.responses.count_tokens.token_counter.openai_count_tokens_handler.handle_count_tokens_request",
new_callable=AsyncMock,
return_value={"input_tokens": 25},
):
result = asyncio.run(
litellm.acount_tokens(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "Hello"}],
system="You are a helpful assistant.",
api_key="sk-test-key",
)
)
assert result.total_tokens == 25
def test_acount_tokens_api_error_falls_back():
"""Test that API errors in token counting return error response."""
from litellm.llms.openai.common_utils import OpenAIError
with patch(
"litellm.llms.openai.responses.count_tokens.token_counter.openai_count_tokens_handler.handle_count_tokens_request",
new_callable=AsyncMock,
side_effect=OpenAIError(status_code=401, message="Invalid API key"),
):
result = asyncio.run(
litellm.acount_tokens(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "Hello"}],
api_key="sk-bad-key",
)
)
# Should fall back to local tokenizer when provider API errors
assert result.error is False
assert result.tokenizer_type == "local_tokenizer"
assert result.total_tokens > 0
def test_acount_tokens_no_api_key_falls_back(monkeypatch):
"""Test that missing API key falls back to local counting."""
monkeypatch.delenv("OPENAI_API_KEY", raising=False)
result = asyncio.run(
litellm.acount_tokens(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "Hello"}],
)
)
# Should fall back to local tokenizer since no API key
assert result.total_tokens > 0
assert result.tokenizer_type == "local_tokenizer"