Commit graph

37309 commits

Author SHA1 Message Date
Ryan Crabbe
27484c4a41
fix: isolate logs team filter dropdown from root teams state bleed
The Logs view's Team ID filter dropdown was reading `allTeams` from the
root `teams` state in page.tsx, which the Teams page search overwrites
with its filtered subset. Applying a team search on the Teams page made
filtered-out teams disappear from the Logs filter dropdown.

Swap the Team ID filter to use the existing `TeamDropdown` component via
a small `FilterTeamDropdown` wrapper that adapts it to the filter slot's
`FilterOptionCustomComponentProps` contract. The dropdown now drives its
own `useInfiniteTeams` query against `/v2/team/list` with server-side
search and an isolated react-query cache, unreachable from root state.

Rename the now-unused `hookAllTeams` destructure to `allTeams` so the
`KeyInfoView` passthrough receives the hook's unpolluted fetch instead
of the polluted prop, and drop the dead `allTeams` prop from
`SpendLogsTable` and both of its call sites.
2026-04-15 09:52:23 -07:00
Ryan Crabbe
f27bf8e711
fix(ui): pre-select backend default for boolean guardrail provider fields
Boolean fields in the auto-generated guardrail provider form (e.g. Noma
`use_v2`) rendered as empty Selects because the Form.Item only populated
`initialValue` for percentage fields, and the `defaultValue` passed to the
Select child was silently dropped by antd's controlled-component wrapper.
Users could not tell what the backend default was, and the visual ambiguity
made flags like `use_v2` look inoperative even though the save path worked.

Unify `initialValue` to fall back through `fieldValue → field.default_value →
(percentage ? 0.5 : undefined)`, and switch Select.Option values from
"true"/"false" strings to real booleans so the backend default flows through
without stringification.
2026-04-15 09:52:23 -07:00
Sameer Kankute
2a198265ac
Fix docs
Some checks are pending
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Waiting to run
Unit Tests: Security / security (push) Waiting to run
2026-04-15 22:18:15 +05:30
Sameer Kankute
1e8fe18662
Fix docs 2026-04-15 22:14:13 +05:30
Sameer Kankute
3fdd67ff23
Delete docs/my-website/blog/debug_cost_discrepancy/index.md 2026-04-15 21:35:05 +05:30
Sameer Kankute
277be4c50e
Add input + output tokens for anthropic message type 2026-04-15 21:08:34 +05:30
Sameer Kankute
e9cf06f381
revert routig related changes 2026-04-15 17:46:21 +05:30
Sameer Kankute
4139e437d1
Fix docs 2026-04-15 17:42:26 +05:30
Sameer Kankute
c03488c661
docs(advisor): revamp advisor tool docs for cross-provider support
- Clearly separate the two tool formats: advisor_20260301 (native/proxy)
  vs litellm_advisor function tool (chat completions with callback)
- Answer "how do I configure the stronger model?" inline where the user
  first encounters each format
- Add supported providers table with native vs orchestration distinction
- Show AdvisorInterceptionLogger setup is required for litellm_advisor
  function format in SDK usage
- Update proxy config example to use model_list deployment names
- Document provider_specific_fields.advisor_tool_results for chat completions
- Document server_tool_use + advisor_tool_result response blocks for messages
- Add max_uses, cost, and streaming behavior notes
- Remove tool_choice="required" from all examples (causes forced loops)

Made-with: Cursor
2026-04-15 17:38:25 +05:30
Sameer Kankute
1f3fc7b5bb
test(advisor): add tests for cross-provider orchestration and streaming
- Test Anthropic executor + OpenAI advisor via orchestration loop
- Test native Anthropic path preserved for claude-opus-4-6 advisor
- Test tool_choice is stripped from follow-up executor turns
- Test provider_specific_fields contains advisor_tool_results blocks
- Test _advisor_interception_converted_stream stays in litellm_params
- Test wildcard router deployment lookup for order fallback

Made-with: Cursor
2026-04-15 17:38:13 +05:30
Sameer Kankute
79ee78ece0
fix(router): wildcard-aware deployment lookup for order-based fallback
Use get_model_list() instead of _get_all_deployments() so pattern-routed
model groups (e.g. openai/* -> openai/gpt-4.1-mini) are included when
computing the deployment order set for fallback logic.

Made-with: Cursor
2026-04-15 17:38:06 +05:30
Sameer Kankute
28b5e2e450
feat(proxy): initialize AdvisorInterceptionLogger from proxy config
Initialize AdvisorInterceptionLogger via initialize_from_proxy_config()
so default_advisor_model and enabled_providers from
litellm_settings.advisor_interception_params are picked up automatically
when using the proxy.

Add advisor_interception_config.yaml example showing the minimal config
needed to wire an advisor model to the proxy.

Made-with: Cursor
2026-04-15 17:38:01 +05:30
Sameer Kankute
2ca1070694
fix(streaming): propagate advisor interception stream flag in http handler
Propagate _advisor_interception_converted_stream from litellm_params into
model_call_details and convert the agentic non-streaming response back to a
fake stream when the flag is set, mirroring the existing websearch path.

Made-with: Cursor
2026-04-15 17:37:55 +05:30
Sameer Kankute
b7e24e7af4
feat(advisor): cross-provider orchestration loop for chat/completions
- Skip agentic loop for Anthropic executor only when advisor is the
  native-compatible claude-opus-4-6; all other advisor models use the
  LiteLLM orchestration loop regardless of executor provider
- Convert litellm_advisor function tools to provider-native format only
  when executor is Anthropic and advisor is claude-opus-4-6; otherwise
  keep as OpenAI-compatible function tool for the orchestration loop
- Add _is_native_anthropic_advisor_model() to resolve proxy aliases
  before checking native compatibility
- Inject server_tool_use + advisor_tool_result into provider_specific_fields
  of the final ModelResponse to match Anthropic native response structure
- Move _advisor_interception_converted_stream flag into litellm_params
  so it is never forwarded to the upstream LLM provider
- Strip tool_choice from optional_params on follow-up executor turns to
  prevent forced advisor re-invocation loops
- Initialize AdvisorInterceptionLogger with default_advisor_model and
  enabled_providers from proxy config via initialize_from_proxy_config()

Made-with: Cursor
2026-04-15 17:37:47 +05:30
Sameer Kankute
1e72e22ebf
feat(messages-api): cross-provider advisor orchestration for /v1/messages
- Run MessagesInterceptor checks before pre-request hooks so the
  synthetic advisor tool is registered before any tool conversion pass
- AdvisorOrchestrationHandler resolves default_advisor_model from
  litellm.advisor_interception_params when the tool definition omits it
- Sub-calls route through llm_router.acompletion() for proper credential
  resolution across any provider
- Native Anthropic path preserved: if executor is Anthropic and advisor
  is claude-opus-4-6 the request passes through unchanged
- Any other combination (Anthropic executor + non-Opus advisor, or
  non-Anthropic executor) uses the LiteLLM orchestration loop
- Final response always includes server_tool_use and advisor_tool_result
  content blocks matching Anthropic native format
- FakeAnthropicMessagesStreamIterator emits correct SSE events for those
  new block types in streaming responses
- Normalize proxy aliases → canonical Anthropic model IDs before
  native passthrough so tools.0.model is never a proxy alias

Made-with: Cursor
2026-04-15 17:37:23 +05:30
Sameer Kankute
ee60f48a45
feat(advisor): add AdvisorInterceptionConfig type definition
Introduces a typed config class for advisor_interception_params used
under litellm_settings in proxy config.

Made-with: Cursor
2026-04-15 17:37:11 +05:30
Yuneng Jiang
6426bc41f5
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_yj_apr14 2026-04-14 22:40:04 -07:00
shin-berri
4a73e94618
Merge pull request #25747 from BerriAI/main
[Infra] Merge main into litellm_internal_staging
2026-04-14 21:13:12 -07:00
yuneng-jiang
72a461ba4a
Merge pull request #25733 from BerriAI/litellm_guardMainBranch
Some checks are pending
CodeQL / Analyze (actions) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
CodSpeed Benchmarks / benchmarks (push) Waiting to run
Helm unit test / unit-test (push) Waiting to run
Read Version from pyproject.toml / read-version (push) Waiting to run
Scorecard supply-chain security / Scorecard analysis (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Waiting to run
Unit Tests: Security / security (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
[Infra] Guard main to only accept PRs from staging and hotfix branches
2026-04-14 20:54:29 -07:00
harish876
b3c413aefe add a composite index on the model_name, model_id and checked_at key for lookup. 2026-04-15 03:41:52 +00:00
yuneng-jiang
9790a46f69
Merge pull request #25730 from BerriAI/yj_bump_apr14_2
bump: version 1.83.7 → 1.83.8
2026-04-14 20:11:58 -07:00
yuneng-jiang
bdb4f396bb
Merge pull request #25741 from joereyna/fix/test-server-root-path-timeout
fix(ci): increase test-server-root-path timeout to 30m
2026-04-14 19:54:23 -07:00
yuneng-jiang
3284dee734
Merge pull request #25737 from joereyna/fix/remove-missing-mcps-coverage-path
fix: remove non-existent litellm_mcps_tests_coverage from coverage combine
2026-04-14 19:54:13 -07:00
yuneng-jiang
50786007fc
Merge pull request #25736 from BerriAI/docs_visual_guide_for_guardrail_fallbacks
docs update
2026-04-14 19:51:54 -07:00
yuneng-jiang
ffc3a9736d
Merge pull request #25739 from BerriAI/litellm_flakyBedrockGptOssToolCall
[Test] Replace flaky bedrock gpt-oss tool-call live test with request-body mock
2026-04-14 19:49:04 -07:00
joereyna
ccbdaa9187
fix(ci): increase test-server-root-path timeout to 30m 2026-04-14 19:42:10 -07:00
Yuneng Jiang
e2043e11f1
[Test] add request-body mock test for bedrock gpt-oss tool schema
Complements the stubbed-out live integration test by verifying the
outgoing Bedrock Converse request body for GPT-OSS is well-formed when
the caller supplies a tool schema with OpenAI-style metadata
($id, $schema, additionalProperties, strict):
- correct converse URL for bedrock/converse/openai.gpt-oss-20b-1:0
- toolConfig.tools[0].toolSpec has the expected name/description
- inputSchema.json keeps type/properties/required and strips fields
  Bedrock does not accept
2026-04-14 19:36:57 -07:00
Yuneng Jiang
8e44a02a22
[Test] stub flaky bedrock gpt-oss function-calling stream test
GPT-OSS on Bedrock intermittently emits truncated toolUse.input deltas
(e.g. accumulated args of '{"":"'), causing
test_function_calling_with_tool_response to hard-fail on json.loads.
The model flakiness is not a litellm regression: the same base test
passes for Anthropic in the same CI run, and the streaming delta path
at invoke_handler.py has not changed recently.

Follow the existing override pattern in TestBedrockGPTOSS
(test_prompt_caching, test_completion_cost, test_tool_call_no_arguments)
and stub the test to pass. The underlying bedrock converse streaming
tool-call path is already covered by Claude/Nova/Llama Converse suites
in test_bedrock_completion.py and test_bedrock_llama.py, so removing
the live GPT-OSS check loses no unique litellm-side signal.
2026-04-14 19:13:42 -07:00
Yuneng Jiang
d6a69b9c81
[Test] mark bedrock gpt-oss function-calling stream test flaky
Bedrock GPT-OSS occasionally emits truncated toolUse.input deltas
(e.g. accumulated args of '{"":"'), which causes
test_function_calling_with_tool_response to hard-fail on json.loads.
Other overrides in TestBedrockGPTOSS already handle similar
model-side flakiness; apply retries=6 delay=5 scoped to this subclass
so other providers keep strict behavior.
2026-04-14 19:10:55 -07:00
Ishaan Jaffer
84c507bc14
fix(mypy): use explicit None check for cache rate values to satisfy type checker 2026-04-14 19:03:56 -07:00
joereyna
a01cf44c35
fix: remove non-existent litellm_mcps_tests_coverage from coverage combine 2026-04-14 18:59:25 -07:00
Yuneng Jiang
38f8d7a008
Point contributors toward litellm_oss_branch in guard error messages 2026-04-14 18:41:59 -07:00
Yuneng Jiang
ab71d3d700
Also reject PRs from forks, not just non-allowlisted branches 2026-04-14 18:39:54 -07:00
shivam
fd110cd5cf
docs update 2026-04-14 18:33:42 -07:00
Ishaan Jaffer
e20d9df1b6
remove test file 2026-04-14 18:29:15 -07:00
Ishaan Jaffer
e0a988e39a
feat(ui/log-details): pass rawInputTokens, cacheReadTokens, cacheCreationTokens to CostBreakdownViewer from SpendLogs 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
0148effd6e
feat(ui/cost-breakdown): show separate Input / Cache Read / Cache Write line items in cost breakdown drawer 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
6b1dc1156e
fix(ui/usage): subtract cache tokens from Input Tokens summary card to avoid double-counting 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
781fc6311b
feat(cost-calculator): compute and store per-type cache costs in CostBreakdown (cache_read_cost, cache_creation_cost) 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
b5a4c26248
feat(logging): pass cache_read_cost and cache_creation_cost through set_cost_breakdown 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
c84597ecd0
fix(bedrock/converse): capture raw input_tokens as text_tokens before cache inflation in PromptTokensDetailsWrapper 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
5c056cae9f
fix(anthropic): store raw text_tokens in PromptTokensDetailsWrapper before cache inflation 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
c1dcfa70c9
feat(types): add cache_read_cost and cache_creation_cost fields to CostBreakdown TypedDict 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
d805fe7103
test(bedrock): add unit tests for cache token billing with prompt caching
Adds TestBedrockInvokeCacheTokenBilling covering the Bedrock InvokeModel path:
- baseline: no cache tokens, prompt_tokens equals input_tokens
- cache_read: prompt_tokens inflated by design, prompt_tokens_details carries breakdown
- cache_creation: same pattern for write tokens
- cost_calculation_correct_with_cache_read: core billing regression test
- cost_calculation_correct_with_cache_creation: write-rate billing regression test
- back_to_back_requests_cost: full end-to-end scenario (cache write then read)

These lock in the fix from PR #25517 - cache tokens were being double-counted
in AnthropicConfig.calculate_usage causing 10-50x inflated cost on cache reads.
2026-04-14 18:29:12 -07:00
Yuneng Jiang
45d1e1b341
[Infra] Guard main branch with PR source-branch check
Adds a GHA that fails PRs to main unless the head branch is
'litellm_internal_staging' or 'litellm_hotfix_*'. Also fails merge_group
events since merge queue is not in use.
2026-04-14 18:19:14 -07:00
yuneng-jiang
5c1f7d99bf
Merge pull request #25731 from BerriAI/docs_guardrail
fallbacks image
2026-04-14 18:13:12 -07:00
shivam
65ce89dc67
update 2026-04-14 18:02:41 -07:00
yuneng-jiang
ec0953f7b0
Merge pull request #25728 from BerriAI/litellm_fix_together_ai_test_model
[Fix] Test - Together AI: replace deprecated Mixtral with serverless Qwen3.5-9B
2026-04-14 18:01:56 -07:00
shivam
19629004f5
fallbacks image 2026-04-14 17:58:11 -07:00
harish876
d20c70f24c Optimize database query which fetches latest model_id, model_name pairs and dedupes them in memory.
Current fix includes
 - Updates test case
 - Optimized query with docstring. The change leverages deduplication and sorting logic from SQL
 - Added a bench script to differentiate peak memory usage before and after
2026-04-15 00:54:37 +00:00