Add _reset_litellm_http_client_cache autouse fixture (matching
test_vertex_gemma_transformation.py) to flush in_memory_llm_clients_cache
before each test. Without this, a cached real AsyncHTTPHandler from an
earlier test could bypass the class-level mock and cause real HTTP calls.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The _safe_get_request_headers caching uses request.state._cached_headers,
which returns a truthy MagicMock on bare MagicMock() objects instead of
None, breaking content-type detection for form-data tests.
Replace instance-level patch.object(client, "post", side_effect=...) with
class-level patch of AsyncHTTPHandler and AsyncMock to reliably intercept
HTTP calls in CI where real Google credentials are available.
The old approach patched a specific instance's post method and passed
client=client to acompletion(). In CI, the mock wasn't intercepting actual
HTTP calls, causing 401 ACCESS_TOKEN_TYPE_UNSUPPORTED errors. The new
approach patches AsyncHTTPHandler at the class level so any instance
created internally by get_async_httpx_client() is also mocked.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two test files were reloading modules in setup_method/fixtures, which
caused class-reference staleness for subsequent tests in the same worker:
1. test_huggingface_embedding_handler.py reloaded
litellm.llms.custom_httpx.http_handler, creating a new HTTPHandler
class. Subsequent tests (e.g. hosted_vllm embedding) created
client = HTTPHandler() from the new class, but llm_http_handler.py
still held the old class reference. isinstance(client, HTTPHandler)
returned False, so a new unpatched client was used and
client.post was never called.
2. test_vertex_ai_rerank_integration.py reloaded
litellm.llms.vertex_ai.rerank.transformation in setup_method,
creating a new VertexAIRerankConfig class. The transformation test
file's module-level import still referenced the old class, so
@patch('...VertexAIRerankConfig._ensure_access_token') patched the
new class while self.config was an instance of the old class,
leaving the mock unapplied and hitting real Google credentials.
Fix: remove the reload calls. The module-level class references are
stable across tests within a worker; the reloads were solving a problem
that doesn't exist and actively created cross-test contamination.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two independent fixes for test_token_counter.py failures in CI:
1. test_disable_hf_tokenizer_download leaked litellm.disable_hf_tokenizer_download=True
because pytest.MonkeyPatch() was never undone. The setting persisted into the
alphabetically-subsequent test_llama2/3_tokenizer_api_failure tests, causing
_select_tokenizer_helper to short-circuit before calling from_pretrained.
Fix: wrap the test body in try/finally and call monkeypatch.undo().
2. encode() returns a HuggingFace Encoding object when the HF tokenizer loads, but
falls back to returning a plain List[int] (tiktoken) when the model hub is
unreachable. test_encoding_and_decoding called .ids on the result, which raises
AttributeError when the list-based fallback is active.
Fix: normalize encode() to always return List[int] by extracting .ids when present,
and remove the now-unnecessary .ids access in the test.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
`os.environ["LAKERA_API_KEY"]` raises KeyError when the env var is absent,
causing test_active_callbacks to error during fixture setup. Switch to
`os.environ.get()` in both lakera_ai.py and lakera_ai_v2.py so initialization
succeeds without the key (actual API calls will fail separately if key is unset).
Also mock `premium_user=True` in the test fixture so the enterprise
`hide_secrets` guardrail can initialize, matching the test's expectations.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
PrismaClient.__init__ does `from prisma import Prisma` inline, which raises
RuntimeError when the Prisma client hasn't been generated. This caused two
tests to fail in CI with:
Exception: Unable to find Prisma binaries. Please run 'prisma generate' first.
Add an autouse fixture that replaces sys.modules['prisma'] with a MagicMock
for the duration of each test, allowing PrismaClient to be instantiated and
client.db to be overridden with the existing mock objects.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- encode() now always returns List[int] by extracting .ids from HuggingFace
Encoding objects, making the return type consistent regardless of tokenizer backend
- test_encoding_and_decoding: remove .ids access since encode() now returns a list
- test_tokenizers: skip llama2 differentiation assertion when HuggingFace tokenizer
is unavailable (CI without network access falls back to tiktoken)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This test has failed repeatedly in CI with:
'Expected _add_prompt_to_generation_params to have been called once. Called 0 times.'
Root cause: _add_prompt_to_generation_params is only called when _supports_prompt()
returns True. Under cross-test state contamination in CI (parallel workers),
langfuse_sdk_version can be in an unexpected state, causing _supports_prompt() to
return False and silently skip the call (exception swallowed by the outer try/except).
Fixes:
- Use reset_mock(side_effect=True) so setUp's trace side_effect is cleared and the
explicit return_value assignment actually takes effect
- Patch _supports_prompt on the logger instance to always return True, making the
_add_prompt_to_generation_params assertion independent of SDK version state
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fix isinstance() checks failing due to module reload in conftest.py.
The conftest.py fixture reloads the litellm module between test modules,
which causes class references imported at module-level to become stale.
When AsyncHTTPHandler is imported at the top of the file and then litellm
is reloaded by the fixture, the isinstance() check fails because the
returned instance is of the NEW AsyncHTTPHandler class while the test
is checking against the OLD class reference.
Solution: Import AsyncHTTPHandler locally within each test function that
uses isinstance() checks. This ensures we get the fresh class reference
after the module reload.
Fixed tests:
- test_session_reuse_integration
- test_get_async_httpx_client_with_shared_session
- test_get_async_httpx_client_without_shared_session
This resolves intermittent CI failures where parallel test execution
triggers the module reload behavior.
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Replace all asyncio.iscoroutinefunction() calls with inspect.iscoroutinefunction()
to fix Python 3.16 deprecation warnings throughout the entire codebase.
Files updated:
- litellm/litellm_core_utils/logging_utils.py
- litellm/proxy/common_utils/performance_utils.py
- litellm/proxy/management_endpoints/key_management_endpoints.py (2 occurrences)
- litellm/proxy/management_endpoints/ui_sso.py
- litellm/litellm_core_utils/redact_messages.py
- litellm/integrations/custom_guardrail.py
- tests/proxy_unit_tests/test_response_polling_handler.py
This addresses Greptile's feedback about incomplete deprecation fixes.
All instances now use the standard library inspect.iscoroutinefunction()
which is the recommended approach and won't be deprecated.
Critical fix for Greptile feedback: The setup_and_teardown fixture was
missing the autouse=True parameter, causing the module reload logic to
never execute. This would result in test pollution as callbacks would
chain across modules.
Changes:
- Add autouse=True to setup_and_teardown fixture in conftest.py
- Add autouse=True to setup_and_teardown fixture in conftest_improved.py
Note: conftest_improved.py is intentionally kept as a reference
implementation showing the recommended improvements. It demonstrates
better patterns for test isolation that can be adopted later.
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
- Replace asyncio.get_event_loop_policy() with asyncio.new_event_loop()
- Use asyncio.set_event_loop() to set the event loop
- Fixes deprecation warning in Python 3.16
- Updated both conftest.py and conftest_improved.py
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Add infrastructure for JSON-declared providers to support /v1/responses
via `supported_endpoints` field in providers.json. Simplify Perplexity
responses config from 410 to 40 lines by moving cost dict→float parsing
to generic validators in ResponseAPIUsage and Usage.
- Add `supported_endpoints` field to SimpleProviderConfig (default: [])
- Add `supports_responses_api()` to JSONProviderRegistry
- Create OpenAILikeResponsesConfig base class for responses API
- Add `create_responses_config_class()` with class caching
- ProviderConfigManager: Python classes take priority over JSON fallback
- Fix ResponseAPIUsage.cost field_validator to handle dict cost objects
- Fix Usage.__init__ to handle dict cost from chat completions
- Simplify PerplexityResponsesConfig with get_supported_openai_params guard
- Add 20 unit tests including Python-over-JSON priority test
- Rename _is_nova_lite_2_model → _is_nova_2_model to match all nova-2-* variants
- Add bedrock/converse/ routing prefix stripping in model detection
- Fix pre-existing test_get_supported_openai_params_bedrock_converse failure
- Remove thinking_blocks tests from Nova 2 test file (not Nova 2 behavior)
- Add end-to-end request, response, multi-turn, and model detection tests
- Parametrize key tests across both nova-2-lite and nova-2-pro model IDs
* allow filtering by user in global usage
* add server root path test to github actions
* Update .github/workflows/test_server_root_path.yml
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* address greptile review feedback (greploop iteration 1)
- Fix HTTPException swallowed by broad except block in get_user_daily_activity
and get_user_daily_activity_aggregated: re-raise HTTPException before the
generic handler so 403 status codes propagate correctly
- Add status_code assertions in non-admin access tests
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* address greptile review feedback (greploop iteration 2)
- Default user_id to caller's own ID for non-admins instead of 403 when
omitted, preserving backward compatibility for API consumers
- Apply same fix to aggregated endpoint
- Update test to verify defaulting behavior instead of expecting 403
- Add useEffect to sync selectedUserId when auth state settles in
UsagePageView to handle async auth initialization
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fixing syntax
* remove artifacts
* feat: guardrail tracing UI - policy, detection method, match details (#21349)
* feat: add GuardrailTracingDetail TypedDict and tracing fields to StandardLoggingGuardrailInformation
* feat: add policy_template field to Guardrail config TypedDict
* feat: accept GuardrailTracingDetail in base guardrail logging method
* feat: populate tracing fields in content filter guardrail
* test: add tracing fields tests for custom guardrail base class
* test: add tracing fields e2e tests for content filter guardrail
* feat: add guardrail tracing UI - policy badges, match details, timeline
* feat: redesign GuardrailViewer to Guardrails & Policy Compliance layout
Two-column layout with request lifecycle timeline on the left
and compact evaluation detail cards on the right. Header shows
guardrail count, pass/fail status, total overhead, policy info,
and an export button.
* feat: add clickable guardrail link in metrics + show policy names
* feat: add risk_score field to StandardLoggingGuardrailInformation
* feat: compute risk_score in content filter guardrail
* feat: display backend risk_score badge on evaluation cards
* fix: fallback to frontend risk score when backend doesn't provide one
* passing in masster key for api calls
* Fix: Add blog as incident report
* Fix: Add blog as incident report
* remove timeline
* feat(models): add github_copilot/gpt-5.3-codex and github_copilot/claude-opus-4.6-fast (#21316)
Add missing GitHub Copilot model entries for gpt-5.3-codex (GA) and
claude-opus-4.6-fast (Public Preview) to both the root and backup
model pricing JSON files.
* only tests for /ui
* bump: version 1.81.12 → 1.81.13
* Fixing mapped tests
* fixing no_config test
* fixing container tests
* fixing test_basic_openai_responses_api
* Adding bedrock thinking budget tokens to docs
* fixing regen key tests
* fix: add missing OpenAI chat completion params to OPENAI_CHAT_COMPLETION_PARAMS
Add store, prompt_cache_key, prompt_cache_retention, safety_identifier, and verbosity
to OPENAI_CHAT_COMPLETION_PARAMS list.
These params were already in DEFAULT_CHAT_COMPLETION_PARAM_VALUES but missing from
the OPENAI_CHAT_COMPLETION_PARAMS list, causing them to be dropped when passed to
OpenAI-compatible providers.
---------
Co-authored-by: yuneng-jiang <yuneng.jiang@gmail.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>
Co-authored-by: Cesar Garcia <128240629+Chesars@users.noreply.github.com>
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
The OpenAI `store` parameter (used for storing completions for
distillation/evals) was missing from `OPENAI_CHAT_COMPLETION_PARAMS`.
This caused it to be unrecognized by `get_standard_openai_params()` and
the `litellm_proxy` provider config. It also meant that code paths using
this list (rather than `DEFAULT_CHAT_COMPLETION_PARAM_VALUES`) would
treat `store` as a provider-specific parameter and forward it to
non-OpenAI providers like Anthropic, resulting in:
"store: Extra inputs are not permitted"
Fixes#19700