Merge pull request #36500 from BerriAI/litellm_feature_request_user_flow

docs: require a user flow and a stuck-at proof in feature requests
This commit is contained in:
devin-ai-integration[bot] 2026-08-11 11:28:41 -07:00 • committed by GitHub
commit 1d90432975
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
4 changed files with 124 additions and 8 deletions

View file

@ -24,10 +24,53 @@ body:
validations:
required: true
- type: textarea
id: motivation
id: user-flow
attributes:
label: Motivation, pitch
description: Please outline the motivation for the proposal. Is your feature request related to a specific problem? e.g., "I'm working on X and would like Y to be possible". If this is related to another GitHub issue, please link here too.
label: User Flow
description: |
Two ordered lists, "Before this feature (today)" and "After this feature (ideal user flow)", walking the same end user through the same task, written strictly from that user's seat. Every rule below applies.
- Describe the real application and the routes its users actually hit, not a generic scenario. Link any related GitHub issue or provider API docs
- Lead each list with one plain sentence saying where the flow dead-ends today and what it would let them do instead, then number the steps
- Every step is something the user does or observes: the HTTP method and full URL they hit, what they sent, and what visibly came back (status code, error text, the shape of an ID). UI steps name the page URL and what is on screen
- No LiteLLM internals: never name functions, files, DB tables, config classes, hooks, callbacks, or code paths. Ask for the behavior you need, not the implementation you imagine
- Keep the two lists step-for-step identical until they diverge, so the missing capability is obvious
- "Before this feature" is also where you show the workaround you're living with, which is what tells us how badly this is needed
placeholder: |
Before this feature (today): a developer batching nightly summaries has no way to mark those calls as low priority, so they compete with live traffic for the same rate limit
1. They send POST https://litellm-domain/v1/chat/completions for 500 documents in a loop
2. Around document 120 they start getting 429s naming the rpm limit, and their user-facing chat app starts getting them too
3. Their workaround is a hand-rolled sleep between calls, which stretches the batch to 3 hours and still collides at peak
After this feature (ideal user flow): the same batch runs as background work that yields to live traffic
1. The developer sends the same POST with "service_tier": "flex"
2. Batch calls queue behind interactive ones instead of 429ing, and the response comes back with the tier it was served at
3. The live chat app keeps returning 200s throughout the batch
4. https://litellm-domain/ui/?page=logs shows the batch requests tagged with that tier
validations:
required: true
- type: textarea
id: how-far-you-got
attributes:
label: How far you got
description: |
Run as many steps of the "After this feature (ideal user flow)" list as you can against a live proxy you ran yourself (e.g., `litellm --config config.yaml --detailed_debug` on localhost:4000), then paste the commands (e.g., curl) and their full output, ending at the step that dead-ends. Every rule below applies.
- Say plainly what stopped you there, in user terms: the option you passed came back ignored, the response 400'd naming an unsupported field, there is no button on the page for it. This is what proves the feature is genuinely missing rather than undocumented
- No mocks. Where the flow involves a provider call, hit the real provider API, even if it costs real $. `pytest` commands are not enough
- Include the config.yaml (or SDK setup) and env vars the proxy ran with, plus the version or commit you were on. Keep the real values for env vars that aren't sensitive, and redact only the secrets: never paste a real API key, virtual key, database URL, or other credential, here or anywhere else in the issue
- If the provider already supports this, link their API docs and paste a direct call to them succeeding, so we can see the shape LiteLLM should be sending
- For UI asks: include screenshots of the page you got stuck on and its URL. Scrub keys and tokens out of screenshots too (for example, the virtual key is briefly shown in the panel right after you create a virtual key)
placeholder: |
Config / setup the proxy ran with:
Version or commit:
Commands and their full output, up to the step that dead-ends:
What stopped me there:
validations:
required: true
- type: dropdown

View file

@ -597,6 +597,13 @@ def build_issue_prompt(*, title: str, body: str) -> str:
that it does not today).
- Motivation / use case with a concrete example (config, API call,
UI flow, or scenario showing what's blocked today).
- END-TO-END EVIDENCE OF THE DEAD-END (set
`has_dead_end_evidence=true` only when this is present): a video,
a screenshot, or the exact command(s) actually run paired with
their real output, showing the point where the flow stops today.
Mocked or stubbed dependencies do NOT count, and an unfilled
template scaffold (bare headings, empty numbered lists) counts as
absent.
For an issue that is neither a bug report nor a feature request (a
question, support request, or discussion), PASS as long as it has a
@ -610,6 +617,7 @@ def build_issue_prompt(*, title: str, body: str) -> str:
"has_repro": boolean,
"has_expected_vs_actual": boolean,
"has_motivation_example": boolean,
"has_dead_end_evidence": boolean,
"missing": ["plain-english strings naming what is missing"],
"explanation": "1-2 sentence reasoning for the team to skim"
}}
@ -707,6 +715,10 @@ _ISSUE_BUG_LABELS: tuple[tuple[str, str], ...] = (
)
_ISSUE_FEATURE_LABELS: tuple[tuple[str, str], ...] = (
("has_motivation_example", "Motivation and concrete example"),
(
"has_dead_end_evidence",
"End-to-end evidence of the dead-end (video, screenshot, or command + real output)",
),
)
@ -838,8 +850,11 @@ def format_issue_close_comment(verdict: dict) -> str:
"video, a screenshot, or the exact commands you ran with their real output / "
"traceback) plus expected vs. actual behavior. Written steps with no run output, "
"video, or screenshot don't count, and mocked or stubbed runs don't count.\n"
" - For **feature requests**: a concrete description of what should change, plus a "
"use case and example (config / API call / UI flow).\n"
" - For **feature requests**: a concrete description of what should change, a "
"use case and example (config / API call / UI flow), plus end-to-end evidence of "
"the dead-end (a video, a screenshot, or the exact commands you ran with their "
"real output showing where the flow stops today). Mocked or stubbed runs don't "
"count.\n"
"2. Comment `@agent-shin reconsider`. I'll re-run triage and reopen the issue if it "
"now meets the bar. (GitHub doesn't let external authors reopen an issue a maintainer "
"or bot closed, so the comment-based reconsider is the reliable path.)\n"
@ -945,8 +960,10 @@ def format_grace_warning_issue_comment(verdict: dict) -> str:
"screenshot, or the exact commands you ran with their real output / traceback) plus "
"expected vs. actual behavior. Written steps with no run output don't count, and "
"mocked or stubbed runs don't count.\n"
"- For **feature requests**: a concrete description of what should change, plus a use "
"case and example (config / API call / UI flow).\n"
"- For **feature requests**: a concrete description of what should change, a use "
"case and example (config / API call / UI flow), plus end-to-end evidence of the "
"dead-end (a video, a screenshot, or the exact commands you ran with their real "
"output showing where the flow stops today). Mocked or stubbed runs don't count.\n"
"\n"
"**If the issue does get auto-closed in 2 hours**, comment `@agent-shin reconsider` "
"and I'll re-evaluate. If it now meets the bar, I'll reopen the issue.\n"

View file

@ -31,7 +31,7 @@ When creating PRs, don't set base to `main`. `litellm_internal_staging` is the d
When writing a PR body, treat the comments and imperative instructions inside @.github/pull_request_template.md as rules to follow, not just layout. Agent harnesses may strip HTML comments from copies of that file injected into context, so read .github/pull_request_template.md from disk before writing a PR body to make sure you see every comment rule
Same applies for filing bug reports and .github/ISSUE_TEMPLATE/bug_report.yml
Same applies for filing bug reports and feature requests, with .github/ISSUE_TEMPLATE/bug_report.yml and .github/ISSUE_TEMPLATE/feature_request.yml, respectively
If you're resolving a linear ticket, in the "## Linear ticket" section of the PR, say "Resolves LIT-1234", replacing "LIT-1234" with the actual ticket id that you're resolving. If you don't have the ticket id, don't make one up or search for it. Just leave the section blank

View file

@ -207,6 +207,23 @@ class TestCloseCommentText:
assert "end-to-end qa proof" in body.lower()
assert "mock" in body.lower()
def test_issue_recovery_comments_should_name_feature_dead_end_evidence(
self, triage_module
):
# The feature-request pass bar demands end-to-end evidence of the
# dead-end, so the close and grace-warning recovery bullets must ask
# for it too — otherwise a requester follows those exact instructions
# (description + use case only) and fails `reconsider` again with no
# hint of what else was needed.
verdict = {"verdict": "fail", "missing": [], "explanation": ""}
for body in (
triage_module.format_issue_close_comment(verdict),
triage_module.format_grace_warning_issue_comment(verdict),
):
normalized = " ".join(body.split())
assert "end-to-end evidence of the dead-end" in normalized
assert "showing where the flow stops today" in normalized
def test_all_agent_shin_comments_should_use_bullet_train_emoji(self, triage_module):
# The bullet train (🚅) is Agent Shin's symbol, matching the LiteLLM
# logo; the previous wave (👋) was generic and didn't match the bot's
@ -289,6 +306,27 @@ class TestCloseCommentText:
assert "Expected vs. actual behavior" in body
assert "- ✅ End-to-end evidence of the bug" not in body
def test_issue_close_comment_should_credit_feature_dead_end_evidence(
self, triage_module
):
# A feature requester who pasted their dead-end run but skipped the
# motivation must see the evidence credited and only the motivation
# listed as a gap — without a dedicated verdict field the praise
# block could never acknowledge the work they did do.
body = triage_module.format_issue_close_comment(
{
"verdict": "fail",
"kind": "feature",
"has_motivation_example": False,
"has_dead_end_evidence": True,
"missing": ["motivation / use case"],
"explanation": "no use case given",
}
)
assert "What you got right" in body
assert "- ✅ End-to-end evidence of the dead-end" in body
assert "- ✅ Motivation and concrete example" not in body
def test_close_comments_should_use_softer_park_for_later_framing(
self, triage_module
):
@ -678,6 +716,24 @@ class TestBuildPrompts:
assert "unfilled template scaffold" in normalized
assert "counts as absent, not as evidence" in normalized
def test_issue_feature_rubric_requires_evidence_of_the_dead_end(
self, triage_module
):
# The feature form asks the requester to walk the ideal flow against a
# live proxy and paste output up to the step that dead-ends, so the
# judge has to demand that evidence, and must not accept an unedited
# scaffold of bare headings as if it were a real attempt.
prompt = triage_module.build_issue_prompt(title="t", body="x")
normalized = " ".join(prompt.split())
assert "END-TO-END EVIDENCE OF THE DEAD-END" in normalized
assert "showing the point where the flow stops today" in normalized
assert "unfilled template scaffold" in normalized
# The evidence has its own verdict field so feature requesters who
# provided it get credited in "What you got right", exactly like
# `has_repro` credits bug evidence.
assert "`has_dead_end_evidence=true` only when this is present" in normalized
assert '"has_dead_end_evidence": boolean' in normalized
def test_should_not_crash_when_pr_body_contains_curly_braces(self, triage_module):
"""User-supplied content with `{` / `}` must NOT be re-parsed by
`str.format()`. `format` only scans the template literal for