mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-14 23:21:35 +00:00
This is the second-step PR of the Agent Shin rollout: it merges 7 days after the heads-up (#28759 + this branch) and turns the bot on for real. Two things happen on merge: 1. A one-shot enactment sweep runs over every open external PR/issue, driving each through the steady-state triage logic with --close=true. PRs that fixed their description in the grace week get tagged `ready for review`; PRs/issues still failing the rubric get the standard 24h grace warning (or close, if they already had one and 24h elapsed). The sweep uses the same `_agent_shin_actions` dry-run wrappers as the heads-up so a single boolean toggles real vs. log. 2. The four existing triage workflows (triage_pr_with_llm, triage_issue_with_llm, close_low_quality_prs, review_gate, triage_reconsider) flip from "dry-run unless AGENT_SHIN_ENABLED=true" to "live unless AGENT_SHIN_ENABLED=false". The variable becomes a kill switch instead of an opt-in. Default semantics: unset means live. Files ----- .github/scripts/triage_rollout_enact.py — the enactment sweep. * Calls review_gate(close=False, now=current_time) for PRs and triage(close=False) (wrapped in `_fake_now(current_time)`) for issues, then routes the verdict through the matching maybe_* wrapper. Two single-page dispatch tables (_apply_pr_result, _apply_issue_result) make it easy to audit which actions map to which mutations. * Time-travel dry-run: --simulate-future-hours N (default 24+1s when --close is not set) shifts the clock forward N hours so you can preview what the next daily cron will do. Implemented in exactly two places: review_gate's `now=` parameter for PRs, and a `_fake_now` context manager that patches `agent_shin_shared.dt.datetime.now` for issues. The context manager restores the original module on exit (and on exception). * --simulate-now ISO_TS pins the clock to a specific timestamp. * --close forbids the simulate flags so a real run is always at wall-clock time. .github/workflows/triage_rollout_enact.yml — thin wrapper. Fires --close on push to litellm_internal_staging (the enactment merge) and offers a workflow_dispatch with dry_run + simulate_future_hours inputs for safe re-runs. .github/workflows/{triage_pr_with_llm,triage_issue_with_llm, close_low_quality_prs,review_gate,triage_reconsider}.yml — inverted: * `${AGENT_SHIN_ENABLED:-true}` (default live) * Conditional: `= "false"` enters kill-switch branch * OPENAI_API_KEY env: exposed unless `vars.AGENT_SHIN_ENABLED == 'false'` Comments updated to call out the new kill-switch semantics. tests/test_litellm/test_github_triage_workflows.py — split the destructive-gate constant into PER_RUN_GATE_ENV (per-input gates that must still match `= "true"`) and KILL_SWITCH_WORKFLOWS (all five workflows, must match the inverted `= "false"` / `!= "false"` pattern). Updates the kill-switch test wording to describe the new semantics. tests/test_litellm/test_triage_rollout_enact.py — 22 tests covering: * _fake_now patches and restores (including on exception) * Each branch of _apply_pr_result / _apply_issue_result with a recorder that captures the maybe_* call sequence and dry_run flag * _process_one skip-not-open / skip-internal-author * run() end-to-end: dry-run threads dry_run=True through every wrapper, real run threads False, current_time threads through to the per-item evaluators, --kind / --only-numbers restrict scope. Local preview commands ---------------------- # What this script would do right now (no GitHub writes): python3 .github/scripts/triage_rollout_enact.py --repo BerriAI/litellm # What the next daily cron will do (24h+1s in the future): python3 .github/scripts/triage_rollout_enact.py --repo BerriAI/litellm \\ --simulate-future-hours 24 # Pin to a specific moment: python3 .github/scripts/triage_rollout_enact.py --repo BerriAI/litellm \\ --simulate-now '2026-06-02T09:00:00Z' Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| agent_tests | ||
| audio_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| enterprise | ||
| guardrails_tests | ||
| image_gen_tests | ||
| integration | ||
| litellm | ||
| litellm-proxy-extras | ||
| litellm_core_utils | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| old_proxy_tests/tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| scim_tests | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| unified_google_tests | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _flush_vcr_cache.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| eval_swe_bench.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| README.MD | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_config.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_entrypoint.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_organizations.py | ||
| test_otel_thread_leak.py | ||
| test_passthrough_endpoints.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/test_litellm
This folder can only run mock tests.