mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-10 03:28:53 +00:00
Address three Greptile/veria-ai concerns on the @agent-shin reconsider flow: 1. **Reconsider had no dry-run path.** The previous reconsider mode ignored `--close` and always posted comments + reopened on a pass. A local operator running `python triage_with_llm.py --reconsider --pr N` would silently take destructive GitHub actions with no way to preview. Reconsider now honors `close=False` the same way regular triage does and returns `would-reopen` / `would-reconsider-still-failing` for step-summary rendering. 2. **Reconsider could reopen maintainer-closed PRs/issues** (Medium security finding from veria-ai). The workflow only checked that the commenter was authorized — it did NOT check that the most recent close was performed by Agent Shin. A contributor could comment `@agent-shin reconsider` on a PR a maintainer closed for non-rubric reasons (duplicate, security report, design rejection) and have the bot reopen it. Add `was_closed_by_agent_shin()` which inspects the issue events API for the most recent `closed` actor and only permits reopen when that actor matches the configured bot login (default `github-actions[bot]`, overridable via env). Fail-closed on missing events. 3. **No rate-limiting on the reconsider trigger.** Every `@agent-shin reconsider` comment burns CI minutes + an OpenAI API call. Add a 10-minute cooldown via `seconds_since_last_reconsider_verdict()` which greps the issue's comment list for the bot's own verdict marker (`<!-- agent-shin:reconsider-verdict -->`). Inside the window the triage returns `skip-rate-limited` and the LLM never runs. Workflow update: - `triage_reconsider.yml` now passes `--close` only when `AGENT_SHIN_ENABLED=true`, matching the pattern of `triage_pr_with_llm.yml`. The script runs in both states so the verdict still appears in the step summary for QA. Tests: - Add 5 reconsider safety tests: dry-run for pass / fail / linked-issue short-circuit, bot-closed-guard refusal on maintainer close, rate-limit refusal inside the cooldown window, and cooldown-elapsed acceptance. - Add unit tests for `was_closed_by_agent_shin` (bot / maintainer / missing actor / env-override) and `seconds_since_last_reconsider_verdict` (no marker / multiple markers / non-bot comment with marker / bot comment without marker). - Pin the `<!-- agent-shin:reconsider-verdict -->` marker in both reopen and still-failing comments — dropping it would silently break the cooldown. Existing reconsider tests updated to pass `close=True` (the production path now) + stub the new guards via `_stub_reconsider_guards`. 112 tests pass (was 93). Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
130 lines
5.8 KiB
YAML
130 lines
5.8 KiB
YAML
name: Agent Shin — reconsider
|
|
|
|
# Comment-trigger workflow: when the PR/issue author (or an internal
|
|
# collaborator) comments `@agent-shin reconsider` on a CLOSED PR/issue,
|
|
# Agent Shin re-runs LLM-judge triage on the current title+body and:
|
|
#
|
|
# - on PASS: posts a "re-evaluated and reopened" comment + reopens.
|
|
# - on FAIL: posts a "still missing X" comment and leaves it closed,
|
|
# so the contributor can iterate again.
|
|
#
|
|
# This exists because GitHub does NOT let an external (non-write-access)
|
|
# OSS contributor reopen a PR/issue closed by a bot or maintainer. Without
|
|
# this comment trigger, a contributor whose PR Agent Shin auto-closed
|
|
# would have no path back into the review queue except opening a fresh PR
|
|
# (which loses the original PR's history). The bot, on the other hand,
|
|
# has write access via GH_TOKEN and can reopen on their behalf.
|
|
#
|
|
# DRY-RUN BY DEFAULT — gated on `vars.AGENT_SHIN_ENABLED == 'true'` just
|
|
# like the other Agent Shin workflows. The workflow also gates on the
|
|
# commenter being either the PR/issue author or an internal collaborator
|
|
# (OWNER/MEMBER/COLLABORATOR) so random commenters cannot DOS the LLM
|
|
# judge or force a reopen.
|
|
|
|
on:
|
|
issue_comment:
|
|
types: [created]
|
|
|
|
permissions:
|
|
contents: read
|
|
issues: write
|
|
pull-requests: write
|
|
|
|
jobs:
|
|
reconsider:
|
|
if: |
|
|
github.repository == 'BerriAI/litellm'
|
|
&& contains(github.event.comment.body, '@agent-shin reconsider')
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- name: Authorize commenter
|
|
# Only the PR/issue author OR an internal collaborator may trigger
|
|
# a reconsider. Outside random commenters could otherwise spam the
|
|
# phrase to burn LLM budget or, if a fail-open bug were ever
|
|
# introduced, force a reopen on someone else's behalf.
|
|
#
|
|
# We expose the authorization decision as a step output and gate
|
|
# every subsequent (potentially destructive) step on it. A `run:`
|
|
# step with `exit 0` would NOT stop the job — only `if:` gating
|
|
# on a known-true output is safe here.
|
|
id: auth
|
|
env:
|
|
COMMENTER: ${{ github.event.comment.user.login }}
|
|
AUTHOR: ${{ github.event.issue.user.login }}
|
|
ASSOCIATION: ${{ github.event.comment.author_association }}
|
|
run: |
|
|
set -euo pipefail
|
|
if [ "${COMMENTER}" = "${AUTHOR}" ]; then
|
|
echo "::notice::Authorized: commenter is the PR/issue author."
|
|
echo "authorized=true" >> "$GITHUB_OUTPUT"
|
|
exit 0
|
|
fi
|
|
case "${ASSOCIATION}" in
|
|
OWNER|MEMBER|COLLABORATOR)
|
|
echo "::notice::Authorized: commenter is an internal collaborator (${ASSOCIATION})."
|
|
echo "authorized=true" >> "$GITHUB_OUTPUT"
|
|
;;
|
|
*)
|
|
echo "::notice::Commenter '${COMMENTER}' (${ASSOCIATION}) is not authorized to trigger reconsider; skipping subsequent steps."
|
|
echo "authorized=false" >> "$GITHUB_OUTPUT"
|
|
;;
|
|
esac
|
|
|
|
- name: Checkout triage script
|
|
if: steps.auth.outputs.authorized == 'true'
|
|
uses: actions/checkout@08eba0b27e820071cde6df949e0beb9ba4906955 # v4.3.0
|
|
with:
|
|
sparse-checkout: .github/scripts
|
|
persist-credentials: false
|
|
|
|
- name: Set up Python
|
|
if: steps.auth.outputs.authorized == 'true'
|
|
uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065 # v5.6.0
|
|
with:
|
|
python-version: "3.12"
|
|
|
|
- name: Install LLM client
|
|
if: steps.auth.outputs.authorized == 'true'
|
|
run: pip install --no-cache-dir "openai>=1.40.0"
|
|
|
|
- name: Run Agent Shin reconsider
|
|
if: steps.auth.outputs.authorized == 'true'
|
|
env:
|
|
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
|
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
OPENAI_BASE_URL: ${{ vars.OPENAI_BASE_URL }}
|
|
TRIAGE_MODEL: ${{ vars.TRIAGE_MODEL }}
|
|
AGENT_SHIN_ENABLED: ${{ vars.AGENT_SHIN_ENABLED }}
|
|
# `issue_comment` events fire for both issues and PR comments.
|
|
# `issue.pull_request` is set iff this is a PR comment, so we use
|
|
# its presence to decide whether to invoke `--pr N` or `--issue N`.
|
|
IS_PR: ${{ github.event.issue.pull_request != null }}
|
|
NUMBER: ${{ github.event.issue.number }}
|
|
run: |
|
|
set -euo pipefail
|
|
if [ "${IS_PR}" = "true" ]; then
|
|
ARGS=(--repo "${{ github.repository }}" --pr "${NUMBER}" --reconsider)
|
|
else
|
|
ARGS=(--repo "${{ github.repository }}" --issue "${NUMBER}" --reconsider)
|
|
fi
|
|
# Reconsider's destructive actions (post comment + reopen) are
|
|
# gated on `--close`, mirroring the regular triage workflows.
|
|
# When AGENT_SHIN_ENABLED is not the EXACT string "true", we
|
|
# still run the script so its verdict + would-X action lands in
|
|
# the step summary for QA — but without `--close`, the script
|
|
# returns `would-reopen` / `would-reconsider-still-failing`
|
|
# instead of touching GitHub state.
|
|
#
|
|
# Use the positive `= "true"` gate (not `!= "true" -> exit`) so
|
|
# the workflow guardrails in
|
|
# tests/test_litellm/test_github_triage_workflows.py see the
|
|
# canonical fail-safe enable pattern. Unknown values like
|
|
# "True", "yes", "1", or typos fall through to the dry-run
|
|
# branch, which is the safe default.
|
|
if [ "${AGENT_SHIN_ENABLED:-false}" = "true" ]; then
|
|
ARGS+=(--close)
|
|
echo "::notice::Agent Shin reconsider ENABLED — running real triage (close=true)."
|
|
else
|
|
echo "::notice::AGENT_SHIN_ENABLED is not 'true' -> reconsider stays in dry-run (no comment, no reopen)."
|
|
fi
|
|
python3 .github/scripts/triage_with_llm.py "${ARGS[@]}"
|