mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-07 02:59:05 +00:00
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): register typesafe as a provider so Jev deployments load Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(decisions): move provider endpoints under llms and validate proxy bodies Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(decisions): add Cloudflare Clef and Strands Decider backends Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): register decisions routes for managed agents and gateway Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(decisions): use raw regex for cloudflare missing account match Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): avoid cast in Cloudflare response unwrapping Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): default model, evaluation health probe, short Cloudflare names The proxy validates only state and questions, so a request without a model falls through to the configured default model like every other route. Health checks probe evaluation-mode deployments through the Decisions API instead of failing with an unsupported mode, and cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short names resolve a mode and a price. The registry no longer claims typed decisions for a provider with no backend. * fix(decisions): let health_check_params override the evaluation probe Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit the decisions endpoint across providers, limits, health and chaos Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port). The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load. * fix(decisions): send env API keys to a configured api_base Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): register the routes through the lazy feature registry The Decisions router was included at import, ahead of the config and DB pass-through endpoints, so a pass-through configured at /v1/decisions was skipped and answered 400 as an unknown Decisions provider. The routes now register through LAZY_FEATURES, which splices them in after every eager route, so a pass-through at /v1/decisions keeps its route while /decisions still serves natively. The lazy OpenAPI snapshot carries the two paths so the schema shows them before the first call. The audit cells add the env-key egress to a configured api_base, the client api_base opt-in shared with chat, the pass-through precedence on an owned proxy, and the Strands evaluation health check resolved from the cost map. The integration config exports the Perplexity env key the first cell needs. * fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell * fix(proxy): let a config pass-through beat a lazily registered route in eager mode With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at startup, so SafeRouteAdder treated a config pass-through at exactly /v1/decisions as already registered and dropped it. In lazy mode a pass-through created through the API after the first native call was skipped the same way. Routes a lazy feature owns no longer count as registered, and a route added at one of their paths is placed ahead of them, the precedence lazy mode gives a config pass-through when the feature has not loaded yet. --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
305 lines
12 KiB
YAML
305 lines
12 KiB
YAML
name: "Unit Tests"
|
|
|
|
on:
|
|
pull_request:
|
|
branches:
|
|
- main
|
|
- "litellm_**"
|
|
push:
|
|
branches:
|
|
- main
|
|
workflow_dispatch:
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
concurrency:
|
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.sha }}
|
|
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
|
|
|
|
# One caller for every tests/test_litellm shard, replacing the nine thin workflow
|
|
# files that each wrapped a single call to _test-unit-base.yml. Adding a shard is
|
|
# now one matrix entry rather than a new file.
|
|
#
|
|
# `name` is the shard id and nothing else, so each check reports as
|
|
# "<shard> / Run tests" exactly as it did when the shard had its own file. Those
|
|
# strings are the branch ruleset's required contexts, so they are load-bearing:
|
|
# renaming an entry renames a required check and the ruleset stops matching it.
|
|
#
|
|
# Every entry states its timeouts even when they equal the base workflow's
|
|
# defaults. An absent matrix key renders as an empty string, which is not a
|
|
# number, so a partially-specified entry would fail the call rather than fall
|
|
# back to the default.
|
|
#
|
|
# tests/unit/proxy keeps its own caller (test-unit-proxy-db.yml): it is already
|
|
# a matrix and carries a shard-coverage guard that reads that file by name.
|
|
# Folding it in here is a follow-up, together with generalising that guard into
|
|
# assert_ci_coverage.py.
|
|
#
|
|
# `unit-flag` names the `.circleci/tests.yml` job that now runs part of the
|
|
# shard under the same Codecov flag. That pipeline is manual-only while the
|
|
# tests migrate, so the shard also runs those files on every event.
|
|
jobs:
|
|
unit:
|
|
name: ${{ matrix.shard }}
|
|
permissions:
|
|
contents: read
|
|
id-token: write
|
|
pull-requests: write
|
|
strategy:
|
|
fail-fast: false
|
|
matrix:
|
|
include:
|
|
- shard: mcp-integration
|
|
artifact-name: mcp-integration
|
|
test-path: "tests/mcp_tests"
|
|
unit-flag: mcp-integration
|
|
workers: 2
|
|
reruns: 0
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: core-utils
|
|
artifact-name: core-utils
|
|
test-path: tests/unit/decisions
|
|
unit-flag: core-utils
|
|
workers: 2
|
|
reruns: 1
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: enterprise-routing
|
|
artifact-name: enterprise-routing
|
|
test-path: ""
|
|
unit-flag: enterprise-routing
|
|
workers: 2
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: integrations
|
|
artifact-name: integrations
|
|
test-path: >-
|
|
tests/test_litellm/integrations
|
|
tests/test_litellm/tracing
|
|
unit-flag: integrations
|
|
workers: 2
|
|
reruns: 3
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: Vertex AI
|
|
artifact-name: llm-vertex-ai
|
|
test-path: ""
|
|
unit-flag: llm-vertex-ai
|
|
workers: 1
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: All Other Providers
|
|
artifact-name: llm-other-providers
|
|
test-path: ""
|
|
unit-flag: llm-other-providers
|
|
workers: 2
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: misc
|
|
artifact-name: misc
|
|
test-path: >-
|
|
tests/test_litellm/test_*.py
|
|
unit-flag: misc
|
|
workers: 2
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: proxy-auth
|
|
artifact-name: proxy-auth
|
|
test-path: >-
|
|
tests/unit/proxy/auth
|
|
tests/unit/proxy/hooks
|
|
tests/unit/proxy/policy_engine
|
|
tests/unit/proxy/client
|
|
--ignore=tests/unit/proxy/auth/test_auth_checks.py
|
|
--ignore=tests/unit/proxy/auth/test_user_api_key_auth.py
|
|
--ignore=tests/unit/proxy/auth/test_default_end_user_budget_simple.py
|
|
--ignore=tests/unit/proxy/auth/test_jwt.py
|
|
--ignore=tests/unit/proxy/auth/test_models_fallback_endpoint.py
|
|
--ignore=tests/unit/proxy/auth/test_multipart_bypass_repro.py
|
|
--ignore=tests/unit/proxy/auth/test_proxy_routes.py
|
|
--ignore=tests/unit/proxy/hooks/test_banned_keyword_list.py
|
|
--ignore=tests/unit/proxy/hooks/test_unit_test_max_model_budget_limiter.py
|
|
workers: 2
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: proxy-endpoints
|
|
artifact-name: proxy-endpoints
|
|
test-path: >-
|
|
tests/unit/proxy/analytics_endpoints
|
|
tests/unit/proxy/decisions_endpoints
|
|
tests/unit/proxy/management_endpoints
|
|
tests/unit/proxy/list_api
|
|
tests/unit/proxy/memory
|
|
tests/unit/proxy/guardrails
|
|
tests/unit/proxy/management_helpers
|
|
--ignore=tests/unit/proxy/management_endpoints/test_jwt_key_mapping.py
|
|
--ignore=tests/unit/proxy/management_endpoints/test_key_generate_prisma.py
|
|
--ignore=tests/unit/proxy/management_endpoints/test_roi_calculator_endpoints.py
|
|
--ignore=tests/unit/proxy/management_helpers/test_audit_logs_proxy.py
|
|
--ignore=tests/unit/proxy/google_endpoints/test_gemini_agents_endpoints.py
|
|
--ignore=tests/unit/proxy/google_endpoints/test_google_endpoint_routing.py
|
|
--ignore=tests/unit/proxy/google_endpoints/test_google_gemini_proxy_request.py
|
|
--ignore=tests/unit/proxy/public_endpoints/test_blog_posts_endpoint.py
|
|
tests/unit/proxy/anthropic_endpoints
|
|
tests/unit/proxy/google_endpoints
|
|
tests/unit/proxy/openai_files_endpoint
|
|
tests/unit/proxy/batches_endpoints
|
|
tests/unit/proxy/container_endpoints
|
|
tests/unit/proxy/fine_tuning_endpoints
|
|
tests/unit/proxy/vector_store_files_endpoints
|
|
tests/unit/proxy/video_endpoints
|
|
tests/unit/proxy/response_api_endpoints
|
|
tests/unit/proxy/image_endpoints
|
|
tests/unit/proxy/ocr_endpoints
|
|
tests/unit/proxy/search_endpoints
|
|
tests/unit/proxy/vector_store_endpoints
|
|
tests/unit/proxy/agent_endpoints
|
|
tests/unit/proxy/a2a
|
|
tests/unit/proxy/credential_endpoints
|
|
tests/unit/proxy/discovery_endpoints
|
|
tests/unit/proxy/health_endpoints
|
|
tests/unit/proxy/shutdown
|
|
tests/unit/proxy/public_endpoints
|
|
tests/unit/proxy/prompts
|
|
tests/unit/proxy/rag_endpoints
|
|
tests/unit/proxy/rerank_endpoints
|
|
tests/unit/proxy/realtime_endpoints
|
|
tests/unit/proxy/ui_crud_endpoints
|
|
tests/unit/proxy/config_resolvers
|
|
tests/unit/proxy/utils
|
|
workers: 4
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: proxy-server
|
|
artifact-name: proxy-server
|
|
test-path: "tests/unit/proxy/proxy_server"
|
|
workers: 4
|
|
reruns: 2
|
|
timeout-minutes: 60
|
|
job-timeout-minutes: 100
|
|
|
|
- shard: proxy-infra
|
|
artifact-name: proxy-infra
|
|
test-path: >-
|
|
tests/unit/proxy/db
|
|
--ignore=tests/unit/proxy/db/db_transaction_queue/test_e2e_pod_lock_manager.py
|
|
--ignore=tests/unit/proxy/db/test_update_daily_tag_spend.py
|
|
tests/unit/proxy/middleware
|
|
--ignore=tests/unit/proxy/middleware/test_request_size_limit_middleware.py
|
|
tests/unit/proxy/spend_tracking
|
|
--ignore=tests/unit/proxy/spend_tracking/test_search_api_logging.py
|
|
tests/unit/proxy/pass_through_endpoints
|
|
tests/unit/proxy/_experimental
|
|
--ignore=tests/unit/proxy/_experimental/mcp_server
|
|
tests/unit/proxy/experimental
|
|
tests/unit/proxy/common_utils
|
|
--ignore=tests/unit/proxy/common_utils/test_cache_aware_routing.py
|
|
--ignore=tests/unit/proxy/common_utils/test_check_batch_cost.py
|
|
--ignore=tests/unit/proxy/common_utils/test_check_responses_cost.py
|
|
--ignore=tests/unit/proxy/common_utils/test_proxy_encrypt_decrypt.py
|
|
--ignore=tests/unit/proxy/common_utils/test_realtime_cache.py
|
|
tests/unit/proxy/enterprise_billing
|
|
tests/unit/proxy/types_utils
|
|
tests/unit/proxy/logging_endpoints
|
|
unit-flag: proxy-infra
|
|
workers: 4
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: proxy-infra-root
|
|
artifact-name: proxy-infra-root
|
|
test-path: >-
|
|
tests/unit/proxy/test_*.py
|
|
--ignore=tests/unit/proxy/test_aproxy_startup.py
|
|
--ignore=tests/unit/proxy/test_credential_slot_registry.py
|
|
--ignore=tests/unit/proxy/test_custom_callback_input.py
|
|
--ignore=tests/unit/proxy/test_custom_logger_s3_gcs.py
|
|
--ignore=tests/unit/proxy/test_custom_tokenizer_bug.py
|
|
--ignore=tests/unit/proxy/test_db_schema_changes.py
|
|
--ignore=tests/unit/proxy/test_deprecated_key_grace_period.py
|
|
--ignore=tests/unit/proxy/test_get_favicon.py
|
|
--ignore=tests/unit/proxy/test_get_image.py
|
|
--ignore=tests/unit/proxy/test_prisma_client_backoff_retry.py
|
|
--ignore=tests/unit/proxy/test_prompt_test_endpoint.py
|
|
--ignore=tests/unit/proxy/test_proxy_config_unit_test.py
|
|
--ignore=tests/unit/proxy/test_proxy_custom_auth.py
|
|
--ignore=tests/unit/proxy/test_proxy_reject_logging.py
|
|
--ignore=tests/unit/proxy/test_proxy_server.py
|
|
--ignore=tests/unit/proxy/test_proxy_setting_guardrails.py
|
|
--ignore=tests/unit/proxy/test_proxy_token_counter.py
|
|
--ignore=tests/unit/proxy/test_proxy_utils.py
|
|
--ignore=tests/unit/proxy/test_reducto_ocr_route.py
|
|
--ignore=tests/unit/proxy/test_response_polling_pre_call_checks.py
|
|
--ignore=tests/unit/proxy/test_server_root_path.py
|
|
--ignore=tests/unit/proxy/test_ui_path_detection.py
|
|
--ignore=tests/unit/proxy/test_unit_test_proxy_hooks.py
|
|
--ignore=tests/unit/proxy/test_update_spend.py
|
|
--ignore=tests/unit/proxy/test_zero_cost_model_budget_bypass.py
|
|
workers: 4
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: caching-local
|
|
artifact-name: caching-local
|
|
test-path: ""
|
|
unit-flag: caching-local
|
|
workers: 2
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: proxy-extras
|
|
artifact-name: proxy-extras
|
|
test-path: ""
|
|
unit-flag: proxy-extras
|
|
workers: 2
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: enterprise-package
|
|
artifact-name: enterprise-package
|
|
test-path: ""
|
|
unit-flag: enterprise-package
|
|
workers: 4
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
|
|
- shard: responses-caching-types
|
|
artifact-name: responses-caching-types
|
|
test-path: ""
|
|
unit-flag: responses-caching-types
|
|
workers: 2
|
|
reruns: 2
|
|
timeout-minutes: 20
|
|
job-timeout-minutes: 60
|
|
uses: ./.github/workflows/_test-unit-base.yml
|
|
with:
|
|
test-path: ${{ matrix.test-path }}
|
|
unit-flag: ${{ matrix.unit-flag || '' }}
|
|
workers: ${{ matrix.workers }}
|
|
reruns: ${{ matrix.reruns }}
|
|
timeout-minutes: ${{ matrix.timeout-minutes }}
|
|
job-timeout-minutes: ${{ matrix.job-timeout-minutes }}
|
|
artifact-name: ${{ matrix.artifact-name }}
|
|
legacy-mcp-peer: ${{ matrix.shard == 'mcp-integration' }}
|