litellm/tests
devin-ai-integration[bot] 3eb7e45615
fix(pricing): drop the unpublished cached rate from the Gemini Live preview entries (#42651)
* fix(pricing): correct cached-token fields on realtime cost-map entries

azure/gpt-realtime-2 was the only member of the gpt-realtime-2 family priced
on one side of its cached-audio meter. Azure publishes that meter as
"gpt-realtime-2 Audio cd inp Gl 1M Tokens" at 0.4 per 1M and charges the
same rate for the write that populates the cache and the read that hits it,
so cache_creation_input_audio_token_cost lands at 4e-07, matching
azure/gpt-realtime-2.1, azure/gpt-realtime-2.1-mini and the openai
gpt-realtime-2 entry. No cost path reads that field yet, so this corrects
what get_model_info reports rather than what anything bills.

The gemini Live entries go the other way. Google's Vertex context-caching
page publishes separate supported-model lists for implicit and explicit
caching, and no Live or native-audio model is in either one. Its pricing
page prints N/A in both cached-input columns for every Gemini 2.5 Flash
Live API row, where plain 2.5 Flash and 2.5 Flash-Lite both carry real
cached prices, and the Vertex model card for the family marks context
caching not supported outright. Vertex never reports cachedContentTokenCount
on a Live session either, including for a byte-identical 7,021-token prefix
replayed across sessions minutes apart, which is well past the 2,048-token
minimum the same page sets for the Gemini 2 family.

So the 7.5e-08 on the two preview siblings priced something the provider does
not sell, and supports_prompt_caching on all three claimed a capability the
model does not have. The rate comes out. The flag is set to false rather than
removed, because get_model_info maps an absent key to None, and None is how
this map spells "nobody checked" across the 2,788 entries that omit it, where
false records the vendor's documented no. Both readers of the flag gate on
`is True`, so nothing bills or behaves differently either way.

Only the cached fields change on the two 09-2025 preview entries. Their
source field points at the Gemini API pricing page rather than the Vertex
one, so they describe a different surface with its own published limits, and
their context windows are left alone rather than assumed to match the Vertex
model card that drives the GA entry.

Tests cover all three halves: the family invariant that a cached audio read
implies an equal cached audio write, a cached count on a Live entry leaving
the bill at the fresh-input total instead of adding the old 7.5e-08, and
supports_prompt_caching answering false for all three entries while still
answering true for 2.5 Flash, so the false cannot be a swallowed lookup
error.

* fix(cost): correct gemini-live-2.5-flash-native-audio limits and capabilities

Google's model card for model ID gemini-live-2.5-flash-native-audio gives a
128K context window and 64K maximum output tokens, and marks structured
output, context caching and URL context as not supported. Its modality list
is text in and out, image in, audio in and out, and video in, with no
document input of any kind.

The entry advertised a 1M context window, an off-by-one 65535 output cap, and
three capability flags the vendor marks unsupported. Context caching is the
fourth and is handled in the cached-fields change alongside its two preview
siblings.

Both the bare id and vertex_ai/gemini-live-2.5-flash-native-audio resolve to
this single entry, so the test drives the corrected values through both.

* test(integration): cover live preview cached tokens billed at the fresh rate

Co-authored-by: Marty Sullivan <marty@martysullivan.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost): cite dated sources for Live entry pins and drop restating docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Marty Sullivan <marty@martysullivan.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 20:35:34 -07:00
..
agent_tests fix(a2a): keep the upstream status on card discovery failures and inject the card client in tests 2026-09-16 23:55:57 +00:00
audio_tests
base_sdk_tests fix(mcp): explain missing public client dependencies 2026-09-19 19:52:13 -07:00
basic_proxy_startup_tests
batches_tests
benchmarks feat(tokenizer): preserve Python defaults with opt-in Rust dispatch (#42174) 2026-09-22 04:41:11 +00:00
code_coverage_tests feat(logger): dispatch Python logging through the Rust diagnostics processor (#42616) 2026-09-22 18:44:15 -07:00
documentation_tests fix(docs-test): restore line-anchored table regex in router settings check 2026-09-18 20:10:20 +00:00
e2e test(e2e): tolerate provider-side flakes on five full-suite cells (#42628) 2026-09-22 20:12:14 -07:00
enterprise
guardrails_tests Revert "Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context" 2026-09-19 17:45:15 +00:00
image_gen_tests test: remove fully commented-out test files that collect no tests 2026-09-17 20:05:27 +00:00
integration fix(pricing): drop the unpublished cached rate from the Gemini Live preview entries (#42651) 2026-09-22 20:35:34 -07:00
litellm-proxy-extras Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-18 20:55:05 -07:00
litellm_utils_tests test: point CircleCI-only suites at models still in the cost map (#42617) 2026-09-22 17:28:34 -07:00
llm_responses_api_testing test: fix five tests left stale by #41311, #41337, #39996 and #41310 2026-09-16 17:47:27 -07:00
llm_translation test: point CircleCI-only suites at models still in the cost map (#42617) 2026-09-22 17:28:34 -07:00
load_tests
local_testing fix: repair seven regressions caught by CircleCI on main (#42640) 2026-09-23 02:26:36 +00:00
logging_callback_tests test: point CircleCI-only suites at models still in the cost map (#42617) 2026-09-22 17:28:34 -07:00
mcp_tests fix(mcp): apply post-call rewrites without stale structured output (#41530) 2026-09-22 13:44:25 -07:00
multi_instance_e2e_tests
ocr_tests test(ocr): replace per-provider OCR test classes with a declarative provider x auth x input matrix 2026-09-18 21:27:03 +00:00
openai_endpoints_tests
otel_tests test: fix stale budget-status and bad-database-url assertions 2026-09-21 15:00:54 -07:00
pass_through_tests fix(mcp): preserve legacy behavior on SDK2 and streamline verification 2026-09-18 22:28:31 -07:00
pass_through_unit_tests feat(proxy): add TinyFish Agent API passthrough with per-step billing (#41099) 2026-09-21 21:21:43 -07:00
proxy_admin_ui_tests
proxy_behavior Merge pull request #42026 from BerriAI/litellm_user_jwt_savings 2026-09-21 16:41:16 -07:00
proxy_e2e_anthropic_messages_tests fix(test): run the all-beta-headers bedrock cases on Claude Fable 5.1 2026-09-19 23:33:29 +00:00
proxy_migration_tests test: count a zombie grandchild as gone in the migrate deploy timeout test (#42570) 2026-09-22 14:59:26 -07:00
proxy_security_tests refactor(proxy): rename the local development override to dangerously_permit_weak_or_unset_master_key so the name says exactly what it permits 2026-09-19 18:53:14 -07:00
proxy_unit_tests fix(proxy): keep the in-flight daily spend batch when shutdown cancels the flush (#42593) 2026-09-22 17:58:08 -07:00
router_unit_tests Merge pull request #42283 from BerriAI/litellm_mid_stream_fallback_walks_full_list 2026-09-21 15:14:04 -07:00
rust-python-harness test: drop tests that pin provider-owned cost map values 2026-09-18 03:55:51 +00:00
search_tests test: remove fully commented-out test files that collect no tests 2026-09-17 20:05:27 +00:00
spend_tracking_tests
store_model_in_db_tests
test_gateway
test_litellm fix(pricing): drop the unpublished cached rate from the Gemini Live preview entries (#42651) 2026-09-22 20:35:34 -07:00
test_litellm_rust refactor(rust): align the cache crates with Python and wire every native backend (#42530) 2026-09-22 13:13:02 -07:00
unified_google_tests test(google_genai): move unified_google_tests to gemini-3.5-flash-lite (#42520) 2026-09-22 13:07:37 -07:00
unit fix(bedrock): send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse (#42644) 2026-09-22 20:09:21 -07:00
vector_store_tests
windows_tests
__init__.py
_fake_openai_endpoint_server.py
_flush_vcr_cache.py
_live_test_helpers.py refactor(test): validate cost-map entries into a typed model 2026-09-16 14:32:08 -07:00
_openai_record_replay_proxy.py
_process_helpers.py test: count a zombie grandchild as gone in the migrate deploy timeout test (#42570) 2026-09-22 14:59:26 -07:00
_vcr_conftest_common.py test(vcr): guard leaked cassette patches and make injected-transport embedding tests immune (#42542) 2026-09-22 14:20:18 -07:00
_vcr_redis_persister.py
_wait_helpers.py
_ws_vcr.py
AGENTS.md ci(tests): wire tests/unit into CircleCI and drain legacy unit shards green 2026-09-20 07:05:42 +00:00
capturing_transport.py test(vcr): guard leaked cassette patches and make injected-transport embedding tests immune (#42542) 2026-09-22 14:20:18 -07:00
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_new_vector_store_endpoints.py
test_openai_endpoints.py test: fix five tests left stale by #41311, #41337, #39996 and #41310 2026-09-16 17:47:27 -07:00
test_organizations.py
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_rust_python_harness.py
test_service_logger_otel.py
test_spend_logs.py
test_team.py
test_team_logging.py
test_team_members.py
test_users.py

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.