mirror of
https://github.com/BerriAI/litellm.git
synced 2026-08-28 05:25:59 +00:00
* feat(logging): add async_post_call_failure_deployment_hook CustomLogger already has async_pre_call_deployment_hook and async_post_call_success_deployment_hook, both firing once per real deployment attempt from wrapper_async since the router re-enters that wrapper fresh on every retry and fallback step. There was no failure-side counterpart; the only failure signal, async_log_failure_event, fires once per logical client request behind a dedup gate, so fallback chain attempts 2+ were invisible to callbacks needing per-deployment-attempt granularity. Adds async_post_call_failure_deployment_hook(request_data, exception, call_type) to CustomLogger and a matching dispatcher in utils.py, called from wrapper_async's except block. It needs no dedup coordination since each real attempt naturally re-enters the wrapper once. Unlike its two siblings, the dispatcher wraps each callback call in its own try/except since it runs on the wrapper's own exception path and a broken callback must never mask the exception about to be re-raised to the caller. * feat(logging): pass fallback_depth through to async_post_call_failure_deployment_hook Router already tracks fallback_depth internally on each fallback hop (litellm/router_utils/fallback_event_handlers.py), incrementing it once per target tried, but nothing surfaced it to CustomLogger callbacks. Reads it off request_data in the dispatcher and passes it through as a best-effort int | None keyword: None on the first, pre-fallback attempt or a bare SDK call with no router, 1 on the first fallback hop, 2 on the second, and so on. Verified live against a real multi-hop Router fallback chain before adding the regression tests. * fix(logging): fire async_post_call_failure_deployment_hook on internal calls too The failure hook was gated behind the same not _is_litellm_internal_call check as the request-level dedup-gated failure logging, so a failed internal sub-call (e.g. an emulated file-search step) never reached it, even though its async_pre_call_deployment_hook and async_post_call_success_deployment_hook siblings already fire unconditionally for such calls. * chore: retrigger CI (lint job hit a transient GitHub Actions infra outage on the prior push) * chore: retrigger CI (lint job hit the same GitHub Actions infra outage again) * fix(logging): scope async_post_call_failure_deployment_hook to the actual model call The hook was dispatched from the wrapper's broad outer except, which also catches BudgetExceededError (raised before any deployment attempt), errors from async_pre_call_deployment_hook, and errors raised after a successful model call (post_call_processing, async_post_call_success_deployment_hook, caching). None of those are a deployment attempt failing, so the hook misreported them as one. Scoped the hook to a try/except around the model call itself, so it only fires when that specific call raises, matching its own documented contract. * test: assert the callback actually ran in the failure-hook error-isolation test An upstream test-quality gate (TQ001) flagged this test for asserting nothing, so it could only fail by raising. Track whether the exploding callback actually ran and assert on it, so the test would catch a dispatcher that silently skipped every callback instead of isolating a raising one. * fix(logging): harden async_post_call_failure_deployment_hook against 5 maintainer-verified issues A maintainer's live-proxy A/B review against base found five real problems with the failure hook, all reproduced and fixed: - The dispatcher called overrides with fallback_depth as a required keyword, so an override matching this PR's own earlier 3-arg proof-of-fix example raised TypeError, swallowed at debug level, on every call. Now checks the override's signature once per class and omits the keyword when unsupported. - A callback mutating the exception it receives (e.g. status_code) changed what the real caller got back, since it was the same live object about to be re-raised. Callbacks now receive a same-class snapshot instead. - request_data exposed attempted_targets, the router's own live fallback-walk bookkeeping shared by reference across every hop, so a callback calling .record() on it could make the router skip a deployment it never actually tried. Now excluded from what the hook receives. - The hook's own await sat directly in the model-call except block, so a caller-side cancellation landing mid-await (e.g. asyncio.wait_for) replaced the real deployment exception with CancelledError/ TimeoutError. Now isolated so hook dispatch can never mask the real failure. - The timestamp used for the reported failure duration was captured after the hook ran, so a slow callback inflated async_log_failure_event's duration. Now captured before the hook dispatches. * fix(logging): preserve traceback/cause/context on the failure-hook exception snapshot Bugbot found a real gap in the previous round's exception-mutation fix: _snapshot_exception_for_hook only copied __dict__ and args, so a callback formatting or inspecting the failure chain saw an empty traceback and lost chained-exception context, even though the live exception still has them. __traceback__/__cause__/__context__ aren't stored in __dict__, so they need copying explicitly. * fix(logging): preserve __suppress_context__ on the failure-hook exception snapshot Setting __cause__ has a documented CPython side effect of implicitly forcing __suppress_context__ to True. Since the previous round's traceback fix set __cause__ before __suppress_context__, a normal implicit-chaining exception (no `raise ... from`, __suppress_context__ naturally False) got its context wrongly suppressed on the snapshot. Now __suppress_context__ is set explicitly, after __cause__, so it always reflects the real exception. * fix(logging): use MappingProxyType for the failure-hook's sanitized request_data A LIT002 budget check (surfaced by rebasing onto a moved base) flagged the dict comprehension building safe_request_data as mutable construction. MappingProxyType is also a strictly better fit here: a genuinely read-only view, not just an immutable-looking dict, matching the intent that callbacks should never be able to mutate what they're handed. --------- Co-authored-by: Deepanshu <deepanshu.lulla@alpha-sense.com> |
||
|---|---|---|
| .. | ||
| agent_tests | ||
| audio_tests | ||
| base_sdk_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| e2e | ||
| enterprise | ||
| guardrails_tests | ||
| image_gen_tests | ||
| integration | ||
| litellm-proxy-extras | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_migration_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| unified_google_tests | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _fake_openai_endpoint_server.py | ||
| _flush_vcr_cache.py | ||
| _live_test_helpers.py | ||
| _openai_record_replay_proxy.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| _wait_helpers.py | ||
| _ws_vcr.py | ||
| eval_swe_bench.py | ||
| fake_openai_endpoint.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| pyrightconfig.json | ||
| README.MD | ||
| test_anthropic_compaction_usage.py | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_organizations.py | ||
| test_otel_thread_leak.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/test_litellm
This folder can only run mock tests.