litellm/tests
mateo-berri 06145a7d71 feat(triage): day-7 enactment sweep + flip AGENT_SHIN_ENABLED to live-by-default
This is the second-step PR of the Agent Shin rollout: it merges 7 days
after the heads-up (#28759 + this branch) and turns the bot on for real.
Two things happen on merge:

1. A one-shot enactment sweep runs over every open external PR/issue,
   driving each through the steady-state triage logic with --close=true.
   PRs that fixed their description in the grace week get tagged
   `ready for review`; PRs/issues still failing the rubric get the
   standard 24h grace warning (or close, if they already had one and
   24h elapsed). The sweep uses the same `_agent_shin_actions` dry-run
   wrappers as the heads-up so a single boolean toggles real vs. log.

2. The four existing triage workflows (triage_pr_with_llm,
   triage_issue_with_llm, close_low_quality_prs, review_gate,
   triage_reconsider) flip from "dry-run unless AGENT_SHIN_ENABLED=true"
   to "live unless AGENT_SHIN_ENABLED=false". The variable becomes a
   kill switch instead of an opt-in. Default semantics: unset means live.

Files
-----
.github/scripts/triage_rollout_enact.py — the enactment sweep.
  * Calls review_gate(close=False, now=current_time) for PRs and
    triage(close=False) (wrapped in `_fake_now(current_time)`) for
    issues, then routes the verdict through the matching maybe_*
    wrapper. Two single-page dispatch tables (_apply_pr_result,
    _apply_issue_result) make it easy to audit which actions map to
    which mutations.
  * Time-travel dry-run: --simulate-future-hours N (default 24+1s
    when --close is not set) shifts the clock forward N hours so you
    can preview what the next daily cron will do. Implemented in
    exactly two places: review_gate's `now=` parameter for PRs, and
    a `_fake_now` context manager that patches
    `agent_shin_shared.dt.datetime.now` for issues. The context
    manager restores the original module on exit (and on exception).
  * --simulate-now ISO_TS pins the clock to a specific timestamp.
  * --close forbids the simulate flags so a real run is always at
    wall-clock time.

.github/workflows/triage_rollout_enact.yml — thin wrapper. Fires
  --close on push to litellm_internal_staging (the enactment merge)
  and offers a workflow_dispatch with dry_run + simulate_future_hours
  inputs for safe re-runs.

.github/workflows/{triage_pr_with_llm,triage_issue_with_llm,
close_low_quality_prs,review_gate,triage_reconsider}.yml — inverted:
  * `${AGENT_SHIN_ENABLED:-true}` (default live)
  * Conditional: `= "false"` enters kill-switch branch
  * OPENAI_API_KEY env: exposed unless `vars.AGENT_SHIN_ENABLED == 'false'`
  Comments updated to call out the new kill-switch semantics.

tests/test_litellm/test_github_triage_workflows.py — split the
  destructive-gate constant into PER_RUN_GATE_ENV (per-input gates
  that must still match `= "true"`) and KILL_SWITCH_WORKFLOWS (all
  five workflows, must match the inverted `= "false"` / `!= "false"`
  pattern). Updates the kill-switch test wording to describe the
  new semantics.

tests/test_litellm/test_triage_rollout_enact.py — 22 tests covering:
  * _fake_now patches and restores (including on exception)
  * Each branch of _apply_pr_result / _apply_issue_result with a
    recorder that captures the maybe_* call sequence and dry_run flag
  * _process_one skip-not-open / skip-internal-author
  * run() end-to-end: dry-run threads dry_run=True through every
    wrapper, real run threads False, current_time threads through to
    the per-item evaluators, --kind / --only-numbers restrict scope.

Local preview commands
----------------------
    # What this script would do right now (no GitHub writes):
    python3 .github/scripts/triage_rollout_enact.py --repo BerriAI/litellm

    # What the next daily cron will do (24h+1s in the future):
    python3 .github/scripts/triage_rollout_enact.py --repo BerriAI/litellm \\
        --simulate-future-hours 24

    # Pin to a specific moment:
    python3 .github/scripts/triage_rollout_enact.py --repo BerriAI/litellm \\
        --simulate-now '2026-06-02T09:00:00Z'

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-25 20:37:52 -07:00
..
agent_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
audio_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
basic_proxy_startup_tests build: migrate packaging, CI, and Docker from Poetry to uv (#25007) 2026-04-09 11:46:23 -07:00
batches_tests chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
benchmarks
code_coverage_tests CI: copy of #25177 (OCI GenAI: embeddings, streaming/reasoning fixes, model catalog) (#28223) 2026-05-23 12:15:41 -07:00
documentation_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
enterprise chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
guardrails_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
image_gen_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
integration CI: copy of #25177 (OCI GenAI: embeddings, streaming/reasoning fixes, model catalog) (#28223) 2026-05-23 12:15:41 -07:00
litellm CI: copy of #25177 (OCI GenAI: embeddings, streaming/reasoning fixes, model catalog) (#28223) 2026-05-23 12:15:41 -07:00
litellm-proxy-extras style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
litellm_core_utils Merge branch 'litellm_internal_staging' into litellm_staging_03_22_2026 2026-04-20 19:56:00 +05:30
litellm_utils_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
llm_responses_api_testing fix(responses): use OpenAI SSEDecoder for Responses API streaming (#28566) 2026-05-22 10:03:36 -07:00
llm_translation Litellm oss staging 04 21 2026 2 (#26569) 2026-05-20 21:25:19 -07:00
load_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
local_testing test(streaming): tolerate Vertex 429 wrapped in MidStreamFallbackError (#28669) 2026-05-22 15:57:29 -07:00
logging_callback_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
mcp_tests feat(mcp): Add tool call and tool list support via UI for Oauth mcps (#28454) 2026-05-22 09:04:04 -07:00
multi_instance_e2e_tests
ocr_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
old_proxy_tests/tests fix: cleanup tests 2026-03-30 16:24:35 -07:00
openai_endpoints_tests chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
otel_tests feat(prometheus): add user_email and user_alias to user budget metrics (#28155) 2026-05-18 16:28:14 -07:00
pass_through_tests chore(deps): bump deps (#28528) 2026-05-22 00:42:21 +00:00
pass_through_unit_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
proxy_admin_ui_tests chore(test): remove dead old Playwright e2e suite (#28632) 2026-05-22 11:29:17 -07:00
proxy_behavior test(proxy): phase-4 payload behavior pinning for tier-2/3 key + team management endpoints (#28681) 2026-05-23 12:16:29 -07:00
proxy_e2e_anthropic_messages_tests chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
proxy_security_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
proxy_unit_tests encrypt callback_vars in key/team metadata at rest (#27141) 2026-05-23 12:15:44 -07:00
router_unit_tests Litellm oss staging 1 (#28337) 2026-05-20 17:27:03 -07:00
scim_tests
search_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
spend_tracking_tests chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
store_model_in_db_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_litellm feat(triage): day-7 enactment sweep + flip AGENT_SHIN_ENABLED to live-by-default 2026-05-25 20:37:52 -07:00
unified_google_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
vector_store_tests fix: drop milvus dbName and partitionNames from MILVUS_OPTIONAL_PARAMS 2026-04-30 11:51:32 -07:00
windows_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
__init__.py
_flush_vcr_cache.py tests(vcr): isolate cassette redis to CASSETTE_REDIS_URL 2026-05-01 12:32:59 -07:00
_vcr_conftest_common.py fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
_vcr_redis_persister.py fix(caching): replay openai/responses bridge cache hits as chat streams (#28158) 2026-05-18 16:27:06 -07:00
eval_swe_bench.py Prompt Compression - add it to the proxy (#25729) 2026-04-20 15:08:00 -07:00
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
README.MD
test_budget_management.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_callbacks_on_proxy.py test(callbacks): harden flaky proxy callback-leak detector (#28195) 2026-05-18 16:39:02 -07:00
test_config.py
test_debug_warning.py
test_default_encoding_non_root.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_end_users.py chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
test_entrypoint.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_health.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_keys.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_litellm_proxy_responses_config.py chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
test_logging.conf
test_models.py test: replace test_add_and_delete_models integration test with mock 2026-03-30 21:30:57 -07:00
test_new_vector_store_endpoints.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_openai_endpoints.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_organizations.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_otel_thread_leak.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_passthrough_endpoints.py
test_presidio_latency.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_proxy_server_non_root.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_ratelimit.py chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
test_resource_cleanup.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_service_logger_otel.py
test_spend_logs.py Litellm oss staging 04 21 2026 2 (#26569) 2026-05-20 21:25:19 -07:00
test_team.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_team_logging.py test: cleanup dead tests 2026-03-28 20:49:02 -07:00
test_team_members.py Litellm oss staging 04 21 2026 2 (#26569) 2026-05-20 21:25:19 -07:00
test_users.py Fix: tag budget reset must drop stale management-cache entry (#27568) 2026-05-10 00:18:55 +00:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.