mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-25 01:02:15 +00:00
* fix(pricing): correct cached-token fields on realtime cost-map entries azure/gpt-realtime-2 was the only member of the gpt-realtime-2 family priced on one side of its cached-audio meter. Azure publishes that meter as "gpt-realtime-2 Audio cd inp Gl 1M Tokens" at 0.4 per 1M and charges the same rate for the write that populates the cache and the read that hits it, so cache_creation_input_audio_token_cost lands at 4e-07, matching azure/gpt-realtime-2.1, azure/gpt-realtime-2.1-mini and the openai gpt-realtime-2 entry. No cost path reads that field yet, so this corrects what get_model_info reports rather than what anything bills. The gemini Live entries go the other way. Google's Vertex context-caching page publishes separate supported-model lists for implicit and explicit caching, and no Live or native-audio model is in either one. Its pricing page prints N/A in both cached-input columns for every Gemini 2.5 Flash Live API row, where plain 2.5 Flash and 2.5 Flash-Lite both carry real cached prices, and the Vertex model card for the family marks context caching not supported outright. Vertex never reports cachedContentTokenCount on a Live session either, including for a byte-identical 7,021-token prefix replayed across sessions minutes apart, which is well past the 2,048-token minimum the same page sets for the Gemini 2 family. So the 7.5e-08 on the two preview siblings priced something the provider does not sell, and supports_prompt_caching on all three claimed a capability the model does not have. The rate comes out. The flag is set to false rather than removed, because get_model_info maps an absent key to None, and None is how this map spells "nobody checked" across the 2,788 entries that omit it, where false records the vendor's documented no. Both readers of the flag gate on `is True`, so nothing bills or behaves differently either way. Only the cached fields change on the two 09-2025 preview entries. Their source field points at the Gemini API pricing page rather than the Vertex one, so they describe a different surface with its own published limits, and their context windows are left alone rather than assumed to match the Vertex model card that drives the GA entry. Tests cover all three halves: the family invariant that a cached audio read implies an equal cached audio write, a cached count on a Live entry leaving the bill at the fresh-input total instead of adding the old 7.5e-08, and supports_prompt_caching answering false for all three entries while still answering true for 2.5 Flash, so the false cannot be a swallowed lookup error. * fix(cost): correct gemini-live-2.5-flash-native-audio limits and capabilities Google's model card for model ID gemini-live-2.5-flash-native-audio gives a 128K context window and 64K maximum output tokens, and marks structured output, context caching and URL context as not supported. Its modality list is text in and out, image in, audio in and out, and video in, with no document input of any kind. The entry advertised a 1M context window, an off-by-one 65535 output cap, and three capability flags the vendor marks unsupported. Context caching is the fourth and is handled in the cached-fields change alongside its two preview siblings. Both the bare id and vertex_ai/gemini-live-2.5-flash-native-audio resolve to this single entry, so the test drives the corrected values through both. * test(integration): cover live preview cached tokens billed at the fresh rate Co-authored-by: Marty Sullivan <marty@martysullivan.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(cost): cite dated sources for Live entry pins and drop restating docstrings Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Marty Sullivan <marty@martysullivan.com> Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| a2a_protocol | ||
| batches | ||
| caching | ||
| chat_completions | ||
| completion_extras | ||
| containers | ||
| endpoints | ||
| enterprise | ||
| expected_fine_tuning_api | ||
| expected_responses_api_request | ||
| experimental_mcp_client | ||
| fixtures/together_ai_sync | ||
| google_genai | ||
| images | ||
| integrations | ||
| interactions | ||
| litellm_core_utils | ||
| llms | ||
| messages | ||
| ocr | ||
| passthrough | ||
| proxy | ||
| rag | ||
| rerank_api | ||
| responses | ||
| router_strategy | ||
| router_utils | ||
| rust_bridge | ||
| secret_managers | ||
| types | ||
| vector_stores | ||
| videos | ||
| __init__.py | ||
| conftest.py | ||
| log.txt | ||
| readme.md | ||
| test_a2a_registry_lookup.py | ||
| test_acompletion_session_reuse_e2e.py | ||
| test_add_deployment_no_master_key.py | ||
| test_aembedding_session_reuse_e2e.py | ||
| test_anthropic_beta_headers_filtering.py | ||
| test_anthropic_skills_transformation.py | ||
| test_assert_ci_coverage.py | ||
| test_assert_workflow_dir_hygiene.py | ||
| test_audio_transcription_rust_bridge.py | ||
| test_auto_update_price_and_context_window_file.py | ||
| test_azure_ad_token_credential_resolution.py | ||
| test_azure_ai_grok_4_3_model_metadata.py | ||
| test_azure_ai_grok_4_6_model_metadata.py | ||
| test_baseten_glm_5_3_model_metadata.py | ||
| test_batch_completion_models_all_responses.py | ||
| test_bedrock_marengo_embed_3_model_metadata.py | ||
| test_budget_ratchet_check.py | ||
| test_chat_ui_responses_session.py | ||
| test_check_licenses.py | ||
| test_check_mcp_operation_boundary.py | ||
| test_check_migrations_no_data_rewrites.py | ||
| test_check_py310_typing_imports.py | ||
| test_check_test_quality.py | ||
| test_check_type_discipline.py | ||
| test_circleci_path_filter.py | ||
| test_circleci_rust_toolchain.py | ||
| test_claude_fable_5_config.py | ||
| test_claude_opus_4_6_config.py | ||
| test_claude_opus_4_8_config.py | ||
| test_claude_opus_5_config.py | ||
| test_claude_sonnet_5_config.py | ||
| test_cloudflare_workers_ai_model_metadata.py | ||
| test_completion_timeout_resolution.py | ||
| test_component_entrypoint.py | ||
| test_compression.py | ||
| test_conftest.py | ||
| test_conftest_isolation.py | ||
| test_constants.py | ||
| test_container_router.py | ||
| test_cost_calculation_log_level.py | ||
| test_cost_calculator.py | ||
| test_cost_map_guard.py | ||
| test_count_tokens_public_api.py | ||
| test_dashscope_image_generation.py | ||
| test_daybreak_model_metadata.py | ||
| test_deepseek_model_metadata.py | ||
| test_default_branch.py | ||
| test_detect_changes.py | ||
| test_dockerfile_apk_repository.py | ||
| test_dockerfile_bedrock_realtime_extra.py | ||
| test_dockerfile_non_root.py | ||
| test_drop_params_env_var.py | ||
| test_e2e_egress_sentinel.py | ||
| test_eager_tiktoken_load.py | ||
| test_env_key_doc_gate.py | ||
| test_exception_exports.py | ||
| test_exception_header_preservation.py | ||
| test_exception_mapping_request_attribute.py | ||
| test_filter_out_litellm_params.py | ||
| test_fireworks_serverless_model_costs.py | ||
| test_gate_slot_lock.py | ||
| test_gemini_3_1_flash_lite_image_pricing.py | ||
| test_gemini_tts_native_audio_pricing.py | ||
| test_get_blog_posts.py | ||
| test_git_hooks.py | ||
| test_gpt_5_4_model_metadata.py | ||
| test_gpt_5_5_model_metadata.py | ||
| test_gpt_image_cost_calculator.py | ||
| test_gpt_realtime_mode.py | ||
| test_groq_streaming_encoding.py | ||
| test_guardrail_exception_status_codes.py | ||
| test_lazy_imports.py | ||
| test_lint_workflow_diff_gates.py | ||
| test_litellm_params_reserved_keys.py | ||
| test_logging.py | ||
| test_lowest_latency_zero_tokens.py | ||
| test_main.py | ||
| test_main_module_header.py | ||
| test_mistral_medium_3_5_model_metadata.py | ||
| test_mistral_small_4_0_model_metadata.py | ||
| test_mistral_zai_glm_5_2_model_metadata.py | ||
| test_model_block_unblock.py | ||
| test_model_cost_aliases.py | ||
| test_model_param_helper.py | ||
| test_model_prices_schema.py | ||
| test_model_response_normalization.py | ||
| test_muse_spark_1_1_model_metadata.py | ||
| test_muse_spark_1_2_model_metadata.py | ||
| test_muse_spark_1_3_model_metadata.py | ||
| test_mutation_report.py | ||
| test_nested_drop_params.py | ||
| test_non_chat_routes_open_llm_spans.py | ||
| test_openai_embedding_encoding_format_default.py | ||
| test_openai_service_tier_long_context_pricing.py | ||
| test_pre_commit_lint.py | ||
| test_prisma_generate_if_needed.py | ||
| test_process_helpers.py | ||
| test_project_alias_tracking.py | ||
| test_project_tags_pydantic.py | ||
| test_proxy_auth.py | ||
| test_rag_openai_ingestion.py | ||
| test_rate_limit_error_unification.py | ||
| test_redact_string_in_error_paths.py | ||
| test_redis.py | ||
| test_redis_credential_provider.py | ||
| test_register_model_custom_pricing.py | ||
| test_register_model_zero_cost_persistence.py | ||
| test_replicate_model_key_format.py | ||
| test_responses_api_bridge_non_stream.py | ||
| test_responses_id_security.py | ||
| test_responses_streaming_container_ownership.py | ||
| test_retrieve_batch_bedrock_dispatch.py | ||
| test_router.py | ||
| test_router_block_helpers.py | ||
| test_router_exception_redaction.py | ||
| test_router_google_genai.py | ||
| test_router_model_cost_isolation.py | ||
| test_router_order_fallback.py | ||
| test_router_per_deployment_num_retries.py | ||
| test_router_redis_init.py | ||
| test_router_retry_backoff_headers.py | ||
| test_router_retry_non_retryable_errors.py | ||
| test_router_retry_policy_update.py | ||
| test_router_silent_experiment.py | ||
| test_router_streaming_fallback_metadata.py | ||
| test_router_weighted_failover.py | ||
| test_ruff_strict_gate.py | ||
| test_sambanova_model_metadata.py | ||
| test_secret_redaction.py | ||
| test_select_ui_test_scope.py | ||
| test_service_logger.py | ||
| test_setup_wizard.py | ||
| test_shared_session_integration.py | ||
| test_ssl_verify_unit.py | ||
| test_stream_chunk_builder_annotations.py | ||
| test_stream_chunk_builder_citations.py | ||
| test_stream_chunk_builder_images.py | ||
| test_streaming_connection_cleanup.py | ||
| test_sync_together_ai_models.py | ||
| test_system_message_format_bug.py | ||
| test_test_quality_gate.py | ||
| test_thinking_enabled.py | ||
| test_together_ai_model_metadata.py | ||
| test_type_check_gate.py | ||
| test_type_discipline_gate.py | ||
| test_typesafe_model_metadata.py | ||
| test_unit_shard_missing_paths.py | ||
| test_unit_shard_per_test_timeout.py | ||
| test_utils.py | ||
| test_utils_module_docstring.py | ||
| test_uuid_helper.py | ||
| test_vcr_safe_body_matcher.py | ||
| test_vertex_ai_xai_grok_prompt_caching_metadata.py | ||
| test_video_generation.py | ||
| test_with_dashboard_node.py | ||
| test_xai_grok_4_3_model_metadata.py | ||
| test_xai_responses_auto_routing.py | ||
Testing for litellm/
This directory 1:1 maps the the litellm/ directory, and can only contain mocked tests.
The point of this is to:
- Increase test coverage of
litellm/ - Make it easy for contributors to add tests for the
litellm/package and easily run tests without needing LLM API keys.
File name conventions
litellm/proxy/test_caching_routes.pymaps tolitellm/proxy/caching_routes.pytest_<filename>.pymaps tolitellm/<filename>.py