Commit graph

42345 commits

Author SHA1 Message Date
nuernber
3db47daefe test(anthropic_messages): add unit tests for _abort_upstream and _enqueue_for_client edge cases
Add test_abort_upstream_logs_warning_when_aclose_raises: verifies that _abort_upstream swallows and logs any exception raised by the upstream's aclose() method instead of propagating it.

Add test_enqueue_for_client_returns_false_when_already_detached: verifies that _enqueue_for_client returns False immediately without touching the queue when client_detached is already set before the call.

Add test_enqueue_for_
2026-08-06 11:17:03 -07:00
nuernber
a85a9e1186 fix(anthropic_messages): strip inline comments, add abort-upstream regression test
Strip net-new inline # blocks from streaming_iterator.py, the unit test file,
and the live-proxy regression test to comply with the no-new-comments rule.

Add test_async_sse_wrapper_aborts_upstream_when_detached_drain_cap_reached:
verifies that when the detached-drain cap is already full, the pump calls
aclose() on the upstream so the provider stops generating and billing
instead of continuing to stream while we record only the partial prefix.

Also fixes LIT001 (bare dict in AsyncIterator union) by replacing dict
with Mapping[str, object] across all three stream-type annotations, and
adds the required LIT003 reason strings to the three noqa: BLE001 directives.
2026-08-06 10:46:54 -07:00
nuernber
321779138e test(env_keys): exclude internal streaming tuning vars from documentation checks
Add ANTHROPIC_MESSAGES_MAX_DETACHED_STREAM_DRAINS and ANTHROPIC_MESSAGES_STREAM_RELAY_QUEUE_MAXSIZE to the excluded set. These are advanced internal infrastructure parameters for streaming/queue management with sensible defaults that most users should not modify.
2026-08-06 10:38:14 -07:00
nuernber
739447fa4b fix(anthropic_messages): bound streaming relay queue and cap detached drains
The relay queue was unbounded, so a client reading a long stream more slowly
than Bedrock produced it let the pump accumulate every pending SSE chunk in
memory, and detached post-disconnect drains had no concurrency bound, so an
authenticated client could open many large streams and read slowly to pin
unbounded worker state.

Bound the relay queue and make the pump apply backpressure while the client is
connected (it blocks on a full queue, racing the disconnect signal), so a slow
reader throttles the upstream read exactly as the old direct yield did. Cap how
many detached drains run at once; over the cap a disconnected pump bills what it
collected instead of draining further. Detached-drain lifetime is otherwise
bounded by the upstream stream/read timeout. Both limits are tunable via env.
2026-08-06 10:38:14 -07:00
nuernber
ce25486702 fix(anthropic_messages): preserve provider error semantics on upstream stream failure
The detached pump previously caught every upstream exception (Bedrock read,
decode, provider-response, or chunk-conversion error) and terminated the
client stream normally, masking the original provider exception and its
status so downstream failure handling never ran.

Now, when the upstream fails while the client is still connected, forward the
original exception through the queue so the client-facing generator re-raises
it and the proxy's failure handling (status code, post_call_failure_hook)
runs unchanged. Only when the client has already disconnected, where there is
no one to propagate to and no failure hook will fire, fall back to salvaging
partial spend from the collected chunks.
2026-08-06 10:38:14 -07:00
nuernber
1b401af716 fix(anthropic_messages): drain upstream in a detached pump so client disconnect doesn't undercount Bedrock spend
On the /v1/messages -> bedrock/ invoke streaming path a client disconnect
raises CancelledError inside the httpx socket read, which unwinds the whole
upstream generator chain before any finally can drain it. Bedrock keeps
generating and billing the full response, so spend tracking logged only the
truncated partial the client drained (output tokens ~1-15 vs the real count)
and undercounted against AWS invocation logs.

Move the upstream read into a detached background task that fully drains the
provider stream to its terminal message_delta/message_stop and bills there.
The client-facing generator only relays chunks off a queue, so a disconnect
tears down the relay but not the pump. A client_detached event stops
enqueueing after disconnect so the queue can't grow unbounded.
2026-08-06 10:38:14 -07:00
yuneng-jiang
48cb89dba7
Merge pull request #36061 from BerriAI/litellm_/blissful-elion-1e2003
fix(proxy): stop resolving the UI session sentinel team on /search_tools/list
2026-08-06 10:04:44 -07:00
Mateo Wang
9e7b05731d
Merge pull request #36054 from BerriAI/litellm_reduce_any_types
refactor(types): cut 653 implicit and explicit Any diagnostics across 11 modules
2026-08-06 09:54:12 -07:00
Mateo Wang
66e6d53931
Merge pull request #36039 from BerriAI/litellm_reload_ledger_test_isolation
test: roll back runtime model registrations between tests
2026-08-06 09:50:39 -07:00
Mateo Wang
41d8cddfd7
Merge pull request #36072 from BerriAI/litellm_ruff_strict_mappingproxy
chore(lint): name MappingProxyType in the mutable-collection fix messages
2026-08-06 09:47:25 -07:00
Mateo Wang
686473b1ab
Merge pull request #36059 from BerriAI/litellm_claude_md_word_budgets
docs: cap all GitHub comments at 15-25 words, curb semicolon splices
2026-08-06 09:47:12 -07:00
Mateo Wang
aae06a5fd1
Merge pull request #36050 from BerriAI/litellm_gate_owned_typecheck_venv
fix(lint): measure the basedpyright budget gate in a gate-owned venv
2026-08-06 09:45:52 -07:00
Praveena Mundolimoole
2d2994c9e9
fix(proxy): yaml store_prompts_in_spend_logs should take precedence over DB cached value (#35769)
When store_model_in_db is true, general_settings are persisted to the
LiteLLM_Config DB table. On subsequent startups and periodic reloads,
_add_general_settings_from_db_config() unconditionally overwrites the
in-memory general_settings with DB-cached values, including
store_prompts_in_spend_logs.

This means a YAML config change (e.g. store_prompts_in_spend_logs: false)
deployed via CI/CD has no effect because the stale DB value (true) always
wins. The admin must manually update via /config/update API after every
deploy, defeating config-as-code.

Fix: track which general_settings keys were explicitly set in YAML at
startup (_yaml_general_settings_keys). During DB config merge, prefer the
YAML value for tracked keys. The DB value is only used as fallback when
YAML does not set the key, preserving the admin UI's ability to change
settings at runtime.

Steps to reproduce:
1. Start proxy with store_model_in_db: true, store_prompts_in_spend_logs: true
2. Change YAML to store_prompts_in_spend_logs: false, restart
3. Send a request, query LiteLLM_SpendLogs - prompts still stored
4. Check LiteLLM_Config table - DB still has true, overriding YAML

Slack thread: https://dataset-jsonhackathon.slack.com/archives/C0ACUS7LM29/p1785835131860139
2026-08-06 09:32:44 -07:00
mateo-berri
ca7453bc69
chore(lint): recompute budget ceilings after merging base
The base branch ratcheted the same limits in 28a277e9, so the conflicting
files were reset to base and the ratchet re-run against the new merge-base
rather than resolved by hand. Each limit is now the base value minus this
branch's own delta, so both ratchets survive: basedpyright -653 across 48
rules, strict ruff -80, LIT -85.
2026-08-06 12:04:38 +00:00
mateo-berri
f9d48bd47c
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_reduce_any_types
# Conflicts:
#	basedpyright-code-budget.json
#	type-discipline-budget.json
2026-08-06 11:48:17 +00:00
mateo-berri
20a94a9100 chore(lint): move MappingProxyType to the dynamic tail of the LIT002 freeze menu
Revert the LIT001 build-clause inserts, phrase the LIT002 menu as
'or (if it really must be dynamic) a MappingProxyType wrapping a dict
literal or comprehension', and fold the two freezing-wrapper exemption
sentences into one that names MappingProxyType beside tuple/frozenset.
2026-08-06 03:58:07 -07:00
Mateo Wang
818e319b53
chore: make it more concise 2026-08-06 03:38:28 -07:00
Mateo Wang
d0712e1a5e
chore: make CLAUDE.md more concise 2026-08-06 03:37:46 -07:00
mateo-berri
d7ca4dc77f chore(lint): say annotate Mapping[...] instead of Mapping alias in typing.Dict ban 2026-08-06 03:34:34 -07:00
mateo-berri
0c1a4b127d chore(lint): name MappingProxyType in the mutable-collection fix messages
LIT001/LIT002 and the typing.Dict ban all steered dict-shaped values to
frozen dataclasses or suppression even though the checker already accepts
MappingProxyType as a freezing wrapper; the messages now name it so the
dict-shaped freeze path is actually discoverable at fix time.
2026-08-06 03:31:40 -07:00
Mateo Wang
b66d4e6965
Merge pull request #35137 from BerriAI/litellm_fix_responses_cost_router_35131
fix(proxy): fetch background responses through the router in CheckResponsesCost
2026-08-06 03:26:36 -07:00
Mateo Wang
972c0d0b04
Merge pull request #35140 from BerriAI/litellm_fix_file_content_placeholder_cost
fix(cost): stop token-pricing the placeholder input on file content calls
2026-08-06 03:04:36 -07:00
Mateo Wang
920e05c484
chore: fix typo 2026-08-06 02:45:44 -07:00
mateo-berri
526f6d793e fix(lint): retire the single-slot base-counts cache
Storing a baseline used to prune every other cache entry, so gate runs in
concurrent worktrees kept evicting each other's baselines and forcing full
recomputes: this bit six times across two nights of benchmarking. The store
now writes alongside existing entries and evicts only the oldest beyond
eight, keyed as before by merge-base and environment fingerprint, so
parallel worktrees' baselines simply coexist
2026-08-06 02:23:52 -07:00
Mateo Wang
729bec69f5
Merge pull request #36014 from BerriAI/litellm_scan_only_tool_results
feat(guardrails): add scan_only_tool_results to scope unified guardrails to tool results
2026-08-06 02:17:29 -07:00
mateo-berri
14d4897e55 fix(guardrails): refuse scan_only_tool_results combos that scan nothing
Prompt Security drops tool and function rows unless check_tool_results
is on, so it now reports scan-only support from that setting and the
registry refuses the pairing at boot. Pairing scan_only_tool_results
with skip_tool_message_in_guardrail excludes every message, so guardrail
initialization now rejects that combination too.
2026-08-06 01:58:06 -07:00
mateo-berri
e824510765 fix(lint): generate the prisma client into the gate-owned venv
prisma resolves its prisma-client-py generator through a plain /bin/sh PATH
lookup, never through the interpreter that ran prisma generate, so the gate's
generate step landed the client in whatever venv the caller had on PATH: the
owned env never received one, every gate run regenerated, the caller's venv
was mutated instead, and any invocation without a venv on PATH (the rewritten
publisher workflow) failed outright

The generate now runs with the target interpreter's bin directory pinned to
the front of the child PATH. The prisma schema joins the environment
fingerprint so clientless counts recorded before this commit can never be
compared against clientful ones, a cold provision announces itself on stderr
instead of sitting silent for two minutes, and the CI gate step reuses the
job's prisma binary cache
2026-08-06 01:54:26 -07:00
Mateo Wang
a0e627f99d
Merge pull request #35468 from elinacse/bugfix/managed-batch-cost-not-logged
fix(batch): track cost for managed batches with no attributable key/u…
2026-08-06 01:51:23 -07:00
mateo-berri
28ff7f3f0b fix(guardrails): scan function-role results and dedupe returned tools
Under scan_only_tool_results, legacy OpenAI function-role messages now count as tool results, and duplicate names among guardrail-returned tools keep only the first occurrence. CustomGuardrail.structured_messages_cover_full_request lets CrowdStrike AIDR declare that its writeback already rebuilds the whole conversation, so handlers install it as-is instead of merging it into the full message list a second time and duplicating out-of-scope rows. Lint budget ceilings ratchet down to match the tree
2026-08-06 00:52:16 -07:00
mateo-berri
0b24a9ab05 merge: litellm_internal_staging into litellm_fix_file_content_placeholder_cost 2026-08-06 00:09:27 -07:00
mateo-berri
7d745521bf fix(guardrails): merge synthesized tools under scan_only_tool_results and reject role-filtered no-op combos at init 2026-08-05 23:46:40 -07:00
mateo-berri
e0c4c7cee0 Merge branch 'litellm_internal_staging' into bugfix/managed-batch-cost-not-logged 2026-08-05 23:42:18 -07:00
Yuneng Jiang
d5a471b2a7
test(proxy): type the search tool test helpers and record lookups with AsyncMock
Replaces the hand-rolled recording double with AsyncMock so the awaited team ids
come from await_args_list instead of a mutated list, and annotates the response
factory now that SearchToolInfoResponse is imported at module level.
2026-08-05 23:40:16 -07:00
Yuneng Jiang
54e9964eb8
fix(proxy): stop resolving the UI session sentinel team on /search_tools/list
Every Admin UI session key is stamped with the reserved team id
`litellm-dashboard`, which never has a row in LiteLLM_TeamTable, so the
team lookup in _filter_visible_search_tools raised 404 and the endpoint
returned 500 for every non-admin dashboard session.

Skip the lookup for that sentinel and scope the caller by its key-level
allowlist alone, matching how MCP and agent permission checks already
treat it. A real team id is still resolved, and a genuine lookup failure
now surfaces with its own status instead of being masked as a 500.
2026-08-05 23:22:11 -07:00
Mateo Wang
3e5dee7317
chore: make it clearer to Claude that GitHub comments must be concise and to use fewer semicolons 2026-08-05 23:21:18 -07:00
Devin AI
55c392bda1 merge litellm_internal_staging 2026-08-06 06:15:17 +00:00
mateo-berri
4f7d1fce3a fix(proxy): fall back to the SDK when a queued response's deployment is missing 2026-08-05 23:13:15 -07:00
mateo-berri
e34b47f682 docs: cap all GitHub comments at 15-25 words, curb semicolon splices
The 15-25 word cap previously applied only to replies/rebuttals to AI PR
review bots; it now covers every GitHub comment (issue comments and PR
discussion comments included). The public-writing punctuation bullet also
gains a warning that a word cap is not a one-sentence cap, so tight budgets
should be met with period splits or conjunctions rather than ";" splices,
at most one ";" per message
2026-08-05 23:11:21 -07:00
Mateo Wang
0acca3e86a
Merge pull request #24548 from mpcusack-altos/fix/bedrock-batch-credential-fields
fix(router): include Bedrock batch/S3 fields and model in deployment credentials
2026-08-05 22:39:39 -07:00
tin-berri
34fc8d2ee7
fix: expired-miss share over all measured turns + cost-optimization tab labels (#36037)
* fix(ui): make the expired-miss stat row a focusable tooltip trigger

* fix: auto-router expired-miss percentage and cost-optimization tab labels

- change expired-miss percentage denominator from return-to-tier misses to
  all measured turns (same_model + first_visit + return_to_tier). when
  auto-routers flip tiers rapidly within TTL, return-to-tier turns become
  hits and disappear from the miss count; the old metric reported only the
  rare failure population. the new metric contextualizes that population as
  a share of overall coverage
- rename usage tab from 'Usage' to 'Overall'
- rename auto-router-usage tab from 'Auto-Router Usage' to 'Auto-Router'
- update component and unit tests to match new semantics
2026-08-05 22:34:55 -07:00
tin-berri
86890654c5
fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day (#36051)
* fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day

* fix(proxy): gate the current-UTC-day extension behind an opt-in param sent by the cost optimization dashboard

* fix(ui): label cost optimization savings dates as UTC days
2026-08-05 22:33:54 -07:00
mateo-berri
4ab7a33d2c
chore(lint): ratchet lint budgets down by what this branch fixed
Lowers the committed ceilings so the headroom shrinks by exactly what was
cleared instead of leaving stale slack for the next change to spend.

basedpyright -653 errors across 48 rules, with reportAny 29204 -> 28842
and reportExplicitAny 9227 -> 9105. Strict ruff -80 violations, led by
ANN401 -59. LIT rules -85, led by LIT001 -76.
2026-08-06 04:50:01 +00:00
Michael Cusack
3d275d97fe fix(router): return model and Bedrock batch fields in deployment credentials
get_deployment_credentials_with_provider dropped s3_region_name,
s3_encryption_key_id, and aws_batch_role_arn because
CredentialLiteLLMParams never declared them, and it never returned the
deployment's model, so proxy batch creation against Bedrock failed with
"LiteLLM doesn't support custom_llm_provider=bedrock for 'create_batch'"
or "AWS IAM role ARN is required" (#25104)

Provider-only file and batch calls keep their no-model contract:
get_team_provider_credentials strips the model key so a provider-scoped
request is not pinned to an arbitrary matching deployment
2026-08-05 21:49:31 -07:00
mateo-berri
e2cb01c87c
fix(types): correct annotations that were false about their runtime values
An adversarial review of the previous commit found annotations that
described what the code wished were true rather than what flows through.
A false annotation is worse than the Any it replaced, since it launders a
wrong assumption past the type checker.

- purview: `_resolve_user_id` claimed every request-body value was a
  Mapping, contradicting `_resolve_trusted_user_id` one method over, which
  types the same argument `Mapping[str, object]`. `_should_block` claimed
  every Graph response value was a sequence of str->str mappings and was
  not assignable from its own producer's return type.
- cato: `_CatoAnalyzeResponse.required_action` was required and
  non-nullable while the API returns null, as seven fixtures in the
  guardrail's own suite assert. `analysis_result` had the same problem.
  The streaming hook narrowed an override parameter below what
  `ProxyLogging` actually passes it.
- marketplace: `_PluginRecord.manifest_json` was `str` against a nullable
  column. Making it honest surfaced a latent crash, covered below.
- ownership: two functions took an attribute Protocol while their own
  bodies branch on `isinstance(response, dict)`, which no Protocol can
  satisfy.
- openapi generator: `paths` claimed every path-item value was an
  operation, though path items also carry `parameters`, `summary` and
  `$ref`.
- custom openapi spec: a TypedDict asserted a shape that the function
  returns raw Pydantic sub-schemas out of. Reverted to Any, which is
  imprecise but not false.

`get_marketplace` did an unguarded `json.loads` on the nullable
`manifest_json` inside an `except json.JSONDecodeError`, which cannot
catch the TypeError a NULL raises, so one NULL row 500s the endpoint. It
now skips the plugin like the file's other two read sites already do, with
a regression test that fails without the guard.

Where honesty cost precision, precision lost. `_should_block` went back to
its original signature entirely: the narrowing needed to type it turned a
fail-closed DLP control fail-open, because the TypeError it used to raise
on a malformed response reached `except Exception` and became a 400.
2026-08-06 04:37:48 +00:00
mateo-berri
fa47c47020 fix(lint): measure the basedpyright budget gate in a gate-owned venv
The gate previously measured whatever environment the caller happened to
have. Locally that is the fat bootstrap venv (--extra proxy pulls in
fastapi-sso, whose type info flips a reportUnnecessaryIsInstance
diagnostic in ui_sso.py), while CI's publisher venv only has the
proxy-dev and e2e-dev groups, so identical trees measured 866 locally vs
865 in CI and every local gate run breached by a phantom +1

scripts/type_check_gate.py now provisions .venv-typecheck itself: a
frozen uv sync of the canonical proxy-dev and e2e-dev groups, the
interpreter pinned to pyrightconfig.json's pythonVersion, plus the
generated Prisma client. Every measurement pass is pinned to that env
with --pythonpath, because basedpyright auto-detects a .venv in the
project root and that auto-detection beats both PATH order and
VIRTUAL_ENV, so the CLI flag is the only pin that actually works. The
dependency-group set is folded into the environment fingerprint, so
artifacts or caches recorded under a different group set never match
and the gate falls back to computing base counts locally instead of
comparing mismatched environments

The publisher workflow drops its own install and prisma steps and lets
the script build the measurement env, and the node heap for the
full-tree pass drops from 12GB to 8GB (peak RSS measured at 5.4GB)
2026-08-05 21:33:24 -07:00
Mateo Wang
ba91768146
Merge pull request #35925 from BerriAI/litellm_tier_aware_reasoning_token_cost
fix(cost): bill reasoning tokens at the service tier output rate
2026-08-05 21:09:55 -07:00
Mateo Wang
b45b4b7300
Merge pull request #35923 from BerriAI/litellm_dated_variant_tier_pricing_sync
fix(pricing): sync flex/priority tier keys to dated OpenAI snapshot variants
2026-08-05 21:09:37 -07:00
tin-berri
7c621b3141
fix(auto-router): accept every reminder marker pair a harness emits (#36029)
* fix(auto-router): accept every reminder marker pair a harness emits

reminder_markers held one (open, close) pair, so a harness that wraps
injected context differently per agent type only got the slice of traffic
using the configured envelope stripped. Every other agent type kept hitting
the original bug: its reminder-only turn never stripped to empty, won
"newest human ask", and the harness blob got classified in place of the
real question, choosing the tier and therefore the spend.

The field now takes a list of ReminderMarkerPair, following the
KeywordTierRule pattern already in this file so each pair validates itself
and errors point at reminder_markers.N.close rather than a bare index.

Blocks from different pairs can nest, which the gap construction could not
handle: resuming the kept text at an inner block's end walks back inside
the enclosing block and leaks its remainder. Running the block ends through
a maximum collapses nested and overlapping spans without a separate merge
pass, and stays linear in block count, which a fold over a growing tuple
of merged spans would not.

A single pair's ends already increase, so the maximum is the identity and
the default path is byte-identical: verified against the shipped function
over 200k generated inputs, and every existing reminder test passes
unchanged. The prior single-pair config shape is rejected loudly at
startup and at /model/new rather than silently stripping nothing.

* docs(auto-router): document reminder_markers in the complexity router README

* chore(ui): regenerate dashboard API types for the reminder_markers shape

---------

Co-authored-by: Abhimanyu Kapur <38531241+akapur99@users.noreply.github.com>
2026-08-05 21:03:36 -07:00
devin-ai-integration[bot]
0bae9708a7
fix(arize_phoenix): lowercase OTLP/gRPC auth metadata key (#34883) 2026-08-05 20:57:50 -07:00
mateo-berri
c2998dea75 fix(guardrails): guard tools write-back under scan_only_tool_results and warn on role-filtered no-op scans 2026-08-05 20:49:15 -07:00