Commit graph

11158 commits

Author SHA1 Message Date
jquinter
acbcfec313
Merge pull request #21422 from BerriAI/fix/lakera-keyerror-missing-api-key
fix(lakera-guardrail): avoid KeyError on missing LAKERA_API_KEY during initialization
2026-02-17 20:06:43 -03:00
jquinter
153af8b901
Merge pull request #21421 from BerriAI/fix/mock-prisma-in-backoff-tests
fix(tests): mock prisma.Prisma in backoff retry tests to avoid 'prisma generate'
2026-02-17 20:04:34 -03:00
jquinter
7b49be1ee1
Merge pull request #21416 from BerriAI/fix/token-counter-hf-fallback
fix(token-counter): normalize encode() return type and handle HF tokenizer fallback
2026-02-17 20:01:42 -03:00
jquinter
6d5f9b5baa
Merge pull request #21419 from BerriAI/fix/langfuse-test-supports-prompt-flakiness
fix(test): prevent flaky failure in test_log_langfuse_v2_handles_null_usage_values
2026-02-17 20:01:03 -03:00
jquinter
afbfbf3cae
Update tests/test_litellm/litellm_core_utils/test_token_counter.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-17 19:52:06 -03:00
Julio Quinteros Pro
f8f2a86ce1 fix(lakera-guardrail): use os.environ.get() to avoid KeyError on missing LAKERA_API_KEY
`os.environ["LAKERA_API_KEY"]` raises KeyError when the env var is absent,
causing test_active_callbacks to error during fixture setup. Switch to
`os.environ.get()` in both lakera_ai.py and lakera_ai_v2.py so initialization
succeeds without the key (actual API calls will fail separately if key is unset).

Also mock `premium_user=True` in the test fixture so the enterprise
`hide_secrets` guardrail can initialize, matching the test's expectations.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-17 19:30:22 -03:00
Julio Quinteros Pro
defc76b32a fix(tests): mock prisma.Prisma in backoff retry tests to avoid 'prisma generate'
PrismaClient.__init__ does `from prisma import Prisma` inline, which raises
RuntimeError when the Prisma client hasn't been generated.  This caused two
tests to fail in CI with:

  Exception: Unable to find Prisma binaries. Please run 'prisma generate' first.

Add an autouse fixture that replaces sys.modules['prisma'] with a MagicMock
for the duration of each test, allowing PrismaClient to be instantiated and
client.db to be overridden with the existing mock objects.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-17 19:29:20 -03:00
Julio Quinteros Pro
16ca7f4f96 fix(token-counter): normalize encode() return type and handle HF tokenizer fallback
- encode() now always returns List[int] by extracting .ids from HuggingFace
  Encoding objects, making the return type consistent regardless of tokenizer backend
- test_encoding_and_decoding: remove .ids access since encode() now returns a list
- test_tokenizers: skip llama2 differentiation assertion when HuggingFace tokenizer
  is unavailable (CI without network access falls back to tiktoken)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-17 19:19:27 -03:00
Julio Quinteros Pro
703f02bf99 fix(test): prevent flaky failure in test_log_langfuse_v2_handles_null_usage_values
This test has failed repeatedly in CI with:
  'Expected _add_prompt_to_generation_params to have been called once. Called 0 times.'

Root cause: _add_prompt_to_generation_params is only called when _supports_prompt()
returns True. Under cross-test state contamination in CI (parallel workers),
langfuse_sdk_version can be in an unexpected state, causing _supports_prompt() to
return False and silently skip the call (exception swallowed by the outer try/except).

Fixes:
- Use reset_mock(side_effect=True) so setUp's trace side_effect is cleared and the
  explicit return_value assignment actually takes effect
- Patch _supports_prompt on the logger instance to always return True, making the
  _add_prompt_to_generation_params assertion independent of SDK version state

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-17 19:18:12 -03:00
Julio Quinteros Pro
26f8e1ac0a fix(tests): resolve test isolation issue in http_handler tests
Fix isinstance() checks failing due to module reload in conftest.py.

The conftest.py fixture reloads the litellm module between test modules,
which causes class references imported at module-level to become stale.
When AsyncHTTPHandler is imported at the top of the file and then litellm
is reloaded by the fixture, the isinstance() check fails because the
returned instance is of the NEW AsyncHTTPHandler class while the test
is checking against the OLD class reference.

Solution: Import AsyncHTTPHandler locally within each test function that
uses isinstance() checks. This ensures we get the fresh class reference
after the module reload.

Fixed tests:
- test_session_reuse_integration
- test_get_async_httpx_client_with_shared_session
- test_get_async_httpx_client_without_shared_session

This resolves intermittent CI failures where parallel test execution
triggers the module reload behavior.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-17 19:01:02 -03:00
Julio Quinteros Pro
890cc08a3a fix: resolve merge conflict and Greptile feedback
- Remove pytest-retry config from pyproject.toml (fixes merge conflict with main)
- Fix asymmetric callback restoration in isolate_litellm_state fixture
  - Now properly saves and restores all callback lists
  - Prevents test pollution from callback state leakage
- Add cache flush after module reload in setup_and_teardown
  - Prevents stale client instances after importlib.reload
- Remove unused imports (curr_dir, Router)

This addresses:
1. Merge conflict in pyproject.toml (CONFLICTING status)
2. Greptile's feedback about asymmetric callback handling
3. Missing cache flush after module reload
4. Code cleanliness (unused variables)
2026-02-17 18:20:44 -03:00
Julio Quinteros Pro
480974e0f9 fix: complete asyncio.iscoroutinefunction deprecation fix across codebase
Replace all asyncio.iscoroutinefunction() calls with inspect.iscoroutinefunction()
to fix Python 3.16 deprecation warnings throughout the entire codebase.

Files updated:
- litellm/litellm_core_utils/logging_utils.py
- litellm/proxy/common_utils/performance_utils.py
- litellm/proxy/management_endpoints/key_management_endpoints.py (2 occurrences)
- litellm/proxy/management_endpoints/ui_sso.py
- litellm/litellm_core_utils/redact_messages.py
- litellm/integrations/custom_guardrail.py
- tests/proxy_unit_tests/test_response_polling_handler.py

This addresses Greptile's feedback about incomplete deprecation fixes.
All instances now use the standard library inspect.iscoroutinefunction()
which is the recommended approach and won't be deprecated.
2026-02-17 18:20:44 -03:00
Julio Quinteros Pro
e37befd5b4 fix: address Greptile feedback - remove duplicates and fix deprecations
- Remove duplicate @pytest.fixture decorator on setup_and_teardown
- Delete conftest_improved.py (duplicate file, pytest only loads conftest.py)
- Remove deprecated event_loop fixture override
- Add asyncio_default_fixture_loop_scope config in pyproject.toml (modern approach)

This fixes pytest-asyncio >=0.22 deprecation warnings while maintaining
session-scoped event loop behavior.
2026-02-17 18:20:43 -03:00
Julio Quinteros Pro
53d9cad8f5 fix(tests): restore autouse=True to setup_and_teardown fixture
Critical fix for Greptile feedback: The setup_and_teardown fixture was
missing the autouse=True parameter, causing the module reload logic to
never execute. This would result in test pollution as callbacks would
chain across modules.

Changes:
- Add autouse=True to setup_and_teardown fixture in conftest.py
- Add autouse=True to setup_and_teardown fixture in conftest_improved.py

Note: conftest_improved.py is intentionally kept as a reference
implementation showing the recommended improvements. It demonstrates
better patterns for test isolation that can be adopted later.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-17 18:20:23 -03:00
Julio Quinteros Pro
3b176a0970 fix(tests): remove asyncio.get_event_loop_policy deprecation warning
- Replace asyncio.get_event_loop_policy() with asyncio.new_event_loop()
- Use asyncio.set_event_loop() to set the event loop
- Fixes deprecation warning in Python 3.16
- Updated both conftest.py and conftest_improved.py

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-17 18:20:23 -03:00
Julio Quinteros Pro
f9ea565b3b fix(tests): improve test isolation in conftest.py
- Move cache flushing to function scope
  - Disable module reload in parallel mode
  - Remove manual event loop creation
2026-02-17 18:20:23 -03:00
Chesars
0c27f20692 feat(openai_like): add Responses API support to JSON provider system
Add infrastructure for JSON-declared providers to support /v1/responses
via `supported_endpoints` field in providers.json. Simplify Perplexity
responses config from 410 to 40 lines by moving cost dict→float parsing
to generic validators in ResponseAPIUsage and Usage.

- Add `supported_endpoints` field to SimpleProviderConfig (default: [])
- Add `supports_responses_api()` to JSONProviderRegistry
- Create OpenAILikeResponsesConfig base class for responses API
- Add `create_responses_config_class()` with class caching
- ProviderConfigManager: Python classes take priority over JSON fallback
- Fix ResponseAPIUsage.cost field_validator to handle dict cost objects
- Fix Usage.__init__ to handle dict cost from chat completions
- Simplify PerplexityResponsesConfig with get_supported_openai_params guard
- Add 20 unit tests including Python-over-JSON priority test
2026-02-17 16:37:38 -03:00
Sameer Kankute
126cf36dc4 move e2e to llm translation 2026-02-17 22:33:13 +05:30
Sameer Kankute
809838042e
Merge pull request #21382 from BerriAI/litellm_vllm_e2e_testing
Add vllm e2e test for embedding
2026-02-17 22:32:42 +05:30
Sameer Kankute
811ffff0b8 move e2e to llm translation 2026-02-17 21:14:44 +05:30
Sameer Kankute
181a1c3a89 Fix test conifg 2026-02-17 21:09:01 +05:30
Sameer Kankute
ec573ee2b0
Merge pull request #21375 from BerriAI/litellm_evals_api
[feat] Add support for Openai Evals API
2026-02-17 21:01:36 +05:30
Sameer Kankute
791cef6d99 fix test_chat_completion 2026-02-17 20:26:28 +05:30
Sameer Kankute
288f7b860c fix test_allow_access_by_email 2026-02-17 20:19:39 +05:30
Sameer Kankute
1ced47c612 fix tests/test_litellm/proxy/_experimental/mcp_server/test_mcp_server.py 2026-02-17 20:14:56 +05:30
Sameer Kankute
550bb621f7 fix llm tests 2026-02-17 20:13:23 +05:30
Sameer Kankute
fe20e66a1d Fix : test_exception_without_scanners 2026-02-17 20:12:02 +05:30
Sameer Kankute
8374b4d939 Fix : test_exception_without_scanners 2026-02-17 20:11:20 +05:30
Sameer Kankute
7a35116148 Fix : test_video_content_handler_uses_get_for_openai 2026-02-17 20:06:08 +05:30
Sameer Kankute
211d6e9d30 Add vllm e2e test for embedding 2026-02-17 19:42:46 +05:30
Sameer Kankute
782b048372 Update tests/test_litellm/llms/openai/evals/test_openai_evals_transformation.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-17 19:30:58 +05:30
Sameer Kankute
32263deb02 Add tests for openai evals 2026-02-17 19:30:58 +05:30
Sameer Kankute
1c2e1148f3
Merge branch 'main' into litellm_oss_staging_02_16_2026 2026-02-17 18:24:56 +05:30
Tomu Hirata
43ba7d0a07 Add test case for Databricks Meta LLaMA 3.1 70B instruct model in content parsing tests
Signed-off-by: Tomu Hirata <tomu.hirata@gmail.com>
2026-02-17 15:36:00 +09:00
yuneng-jiang
8576683f39
Update tests/test_litellm/proxy/management_endpoints/test_key_management_endpoints.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-16 21:09:13 -08:00
yuneng-jiang
ee5120bfc2 fix default duration 2026-02-16 20:54:10 -08:00
Emerson Gomes
d859f0687d
fix(router): avoid alias scan for non-alias get_model_list lookups (#21136)
Co-authored-by: Codex <codex@example.com>
2026-02-16 20:40:24 -08:00
Ryan H
5749ca6c47 feat(bedrock): broaden Nova 2 model detection to support nova-2-pro reasoning
- Rename _is_nova_lite_2_model → _is_nova_2_model to match all nova-2-* variants
- Add bedrock/converse/ routing prefix stripping in model detection
- Fix pre-existing test_get_supported_openai_params_bedrock_converse failure
- Remove thinking_blocks tests from Nova 2 test file (not Nova 2 behavior)
- Add end-to-end request, response, multi-turn, and model detection tests
- Parametrize key tests across both nova-2-lite and nova-2-pro model IDs
2026-02-16 20:35:00 -08:00
shin-bot-litellm
b609f5841b
fix: add missing OpenAI chat completion params to OPENAI_CHAT_COMPLETION_PARAMS (#21360)
* allow filtering by user in global usage

* add server root path test to github actions

* Update .github/workflows/test_server_root_path.yml

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* address greptile review feedback (greploop iteration 1)

- Fix HTTPException swallowed by broad except block in get_user_daily_activity
  and get_user_daily_activity_aggregated: re-raise HTTPException before the
  generic handler so 403 status codes propagate correctly
- Add status_code assertions in non-admin access tests

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* address greptile review feedback (greploop iteration 2)

- Default user_id to caller's own ID for non-admins instead of 403 when
  omitted, preserving backward compatibility for API consumers
- Apply same fix to aggregated endpoint
- Update test to verify defaulting behavior instead of expecting 403
- Add useEffect to sync selectedUserId when auth state settles in
  UsagePageView to handle async auth initialization

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fixing syntax

* remove artifacts

* feat: guardrail tracing UI - policy, detection method, match details (#21349)

* feat: add GuardrailTracingDetail TypedDict and tracing fields to StandardLoggingGuardrailInformation

* feat: add policy_template field to Guardrail config TypedDict

* feat: accept GuardrailTracingDetail in base guardrail logging method

* feat: populate tracing fields in content filter guardrail

* test: add tracing fields tests for custom guardrail base class

* test: add tracing fields e2e tests for content filter guardrail

* feat: add guardrail tracing UI - policy badges, match details, timeline

* feat: redesign GuardrailViewer to Guardrails & Policy Compliance layout

Two-column layout with request lifecycle timeline on the left
and compact evaluation detail cards on the right. Header shows
guardrail count, pass/fail status, total overhead, policy info,
and an export button.

* feat: add clickable guardrail link in metrics + show policy names

* feat: add risk_score field to StandardLoggingGuardrailInformation

* feat: compute risk_score in content filter guardrail

* feat: display backend risk_score badge on evaluation cards

* fix: fallback to frontend risk score when backend doesn't provide one

* passing in masster key for api calls

* Fix: Add blog as incident report

* Fix: Add blog as incident report

* remove timeline

* feat(models): add github_copilot/gpt-5.3-codex and github_copilot/claude-opus-4.6-fast (#21316)

Add missing GitHub Copilot model entries for gpt-5.3-codex (GA) and
claude-opus-4.6-fast (Public Preview) to both the root and backup
model pricing JSON files.

* only tests for /ui

* bump: version 1.81.12 → 1.81.13

* Fixing mapped tests

* fixing no_config test

* fixing container tests

* fixing test_basic_openai_responses_api

* Adding bedrock thinking budget tokens to docs

* fixing regen key tests

* fix: add missing OpenAI chat completion params to OPENAI_CHAT_COMPLETION_PARAMS

Add store, prompt_cache_key, prompt_cache_retention, safety_identifier, and verbosity
to OPENAI_CHAT_COMPLETION_PARAMS list.

These params were already in DEFAULT_CHAT_COMPLETION_PARAM_VALUES but missing from
the OPENAI_CHAT_COMPLETION_PARAMS list, causing them to be dropped when passed to
OpenAI-compatible providers.

---------

Co-authored-by: yuneng-jiang <yuneng.jiang@gmail.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>
Co-authored-by: Cesar Garcia <128240629+Chesars@users.noreply.github.com>
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2026-02-16 20:31:21 -08:00
Emerson Gomes
b67c140938
Fix Bedrock service_tier cost propagation (#21172) 2026-02-16 20:30:10 -08:00
Nick Amabile
4978df8ebd
fix: add store to OPENAI_CHAT_COMPLETION_PARAMS (#21195)
The OpenAI `store` parameter (used for storing completions for
distillation/evals) was missing from `OPENAI_CHAT_COMPLETION_PARAMS`.

This caused it to be unrecognized by `get_standard_openai_params()` and
the `litellm_proxy` provider config. It also meant that code paths using
this list (rather than `DEFAULT_CHAT_COMPLETION_PARAM_VALUES`) would
treat `store` as a provider-specific parameter and forward it to
non-OpenAI providers like Anthropic, resulting in:

    "store: Extra inputs are not permitted"

Fixes #19700
2026-02-16 20:28:34 -08:00
sahukanishka
d184b3cae7
fix: preserve provider_specific_fields from proxy responses (#21153) (#21220)
Co-authored-by: kanishka sahu <kanishkasahu@mercor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-16 20:17:58 -08:00
Emerson Gomes
f162371b93
fix(pod-lock): make release lock compare-and-delete atomic (#21226) 2026-02-16 20:15:46 -08:00
yuneng-jiang
45e6440b0a fixing test_basic_openai_responses_api 2026-02-16 20:14:06 -08:00
yuneng-jiang
23219d9217 fixing container tests 2026-02-16 20:13:55 -08:00
yuneng-jiang
efe84777e5 fixing no_config test 2026-02-16 20:13:45 -08:00
yuneng-jiang
349e3dad55 Fixing mapped tests 2026-02-16 20:13:37 -08:00
Mateusz Szewczyk
72af441159
feat: Add IBM watsonx.ai rerank support (#21303)
* feat: Add IBM watsonx.ai rerank support

* feat: added unit tests

* fix docstring

* added documentataion

* Update litellm/llms/watsonx/rerank/transformation.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Update litellm/rerank_api/main.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Update litellm/llms/watsonx/rerank/transformation.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* update validate_environment signature

* fix ruff check and mypy

* fix CR

* CR fix

* CR fix

---------

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-16 20:12:16 -08:00
Adam Reed
8d50956051
fix(proxy): preserve and forward OAuth Authorization headers through proxy layer (#19912)
PR #21039 fixed OAuth token handling at the LLM layer (Authorization: Bearer
instead of x-api-key), but the proxy layer still strips the Authorization
header in clean_headers() before it reaches the Anthropic code. This breaks
OAuth for proxy users (e.g., Claude Code Max through LiteLLM proxy).

Changes:
- Add is_anthropic_oauth_key() helper to detect OAuth tokens (sk-ant-oat*)
- Preserve OAuth Authorization headers in clean_headers() instead of stripping
- Forward OAuth Authorization via ProviderSpecificHeader in
  add_provider_specific_headers_to_request() so tokens only reach
  Anthropic-compatible providers (anthropic, bedrock, vertex_ai)

Fixes #19618

Co-authored-by: Adam Reed <iamadamreed@users.noreply.github.com>
2026-02-16 20:09:07 -08:00
yuneng-jiang
df5e8d01a6 address greptile review feedback (greploop iteration 2)
- Default user_id to caller's own ID for non-admins instead of 403 when
  omitted, preserving backward compatibility for API consumers
- Apply same fix to aggregated endpoint
- Update test to verify defaulting behavior instead of expecting 403
- Add useEffect to sync selectedUserId when auth state settles in
  UsagePageView to handle async auth initialization

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-02-16 17:37:18 -08:00