mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
Addresses findings from three independent code reviews (Claude Opus 5, GPT-5.6-sol xhigh via codex, Grok 4.6 xhigh via cursor) ahead of proposing this branch upstream. Correctness: * The restored model qualifier was `clinepass/`, a namespace that does not exist. The catalog namespace is `cline-pass/` (hyphenated); `clinepass/` is LiteLLM's own routing prefix, which is stripped before the request is built. This failed silently because the API validates only the *shape* of a model id -- `totallybogus/deepseek-v4-flash` also returns HTTP 200. It is not inert, though: verified against the live API, `cline-pass/deepseek-v4-flash` resolves to `deepseek/deepseek-v4-flash` while any unrecognised namespace falls back to the date-pinned `deepseek/deepseek-v4-flash-0731`. * `_correct_truncated_finish_reason` compared an aggregate `usage.completion_tokens` against a per-choice `max_tokens`. With `n > 1` that relabels naturally-finished choices as truncated: two 60-token choices under a cap of 100 aggregate to 120 and both became `length`. Restricted to single-choice responses, where the inference is sound. * Zero, negative and `bool` caps are now rejected. `bool` subclasses `int`, so `max_tokens=True` was read as a cap of 1 and would relabel everything. Float usage/caps are now accepted rather than silently skipped. * `_unwrap_response_envelope` reached for the private `_request` attribute because `httpx.Response.request` raises `RuntimeError` instead of returning `None`. Ask for the public attribute defensively instead. * `get_models()` is overridden to return an empty catalog. ClinePass has no `/models` endpoint (404), and the inherited OpenAI implementation asked for it at the wrong path. * Dropped the `async_transform_request` override. `BaseLLMHTTPHandler` builds the body with the synchronous `transform_request` on both the sync and async paths, so it was dead code -- the same shape of bug this provider shipped once already. Pinned with a test. Packaging / CI: * Added the `clinepass` entry to `provider_endpoints_support.json` and its backup. Without it `check_provider_folders_documented.py` fails, which is a required Code Quality job -- verified failing before, passing after. * `transformation.py` did not satisfy `ruff format` under the repo's pinned ruff 0.15.3, which the changed-file CI gate runs. Reformatted. * This commit replaces the previous branch tip, which had accidentally swept in ~435 regenerated Next.js artifacts under `litellm/proxy/_experimental/`. The branch is now +883/-0 across 12 files. Honesty note on the truncation fix: re-probing the live API (streaming and non-streaming, caps of 2000 and 4000, both namespaces) could NOT reproduce the `stop`-instead-of-`length` misreport that motivated it. The correction is kept as a conservative safety net and documented as such rather than as a workaround for a currently-observable defect. Streaming is deliberately not covered; see the docstring. Tests: 41 pass (was 32). openai_like + cometapi regressions: 126 passed, 8 skipped. (cherry picked from commit c2dda4fd8b7a4e1874642b25d73656292cb52468) |
||
|---|---|---|
| .. | ||
| _support | ||
| agent_tests | ||
| audio_tests | ||
| base_sdk_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| e2e | ||
| guardrails_tests | ||
| harness_e2e | ||
| image_gen_tests | ||
| integration | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_migration_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| rust-python-harness | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| test_litellm_rust | ||
| unified_google_tests | ||
| unit | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _fake_openai_endpoint_server.py | ||
| _flush_vcr_cache.py | ||
| _live_test_helpers.py | ||
| _openai_record_replay_proxy.py | ||
| _process_helpers.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| _wait_helpers.py | ||
| _ws_vcr.py | ||
| AGENTS.md | ||
| capturing_transport.py | ||
| eval_swe_bench.py | ||
| fake_openai_endpoint.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| pyrightconfig.json | ||
| README.MD | ||
| test_anthropic_compaction_usage.py | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_otel_thread_leak.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_rust_python_harness.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
| white_100x100.png | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/unit
This folder can only run mock tests.