Commit graph

37693 commits

Author SHA1 Message Date
Ishaan Jaffer
b8b22c9844
fix(types/guardrails): add input_type and messages to ApplyGuardrailRequest 2026-04-23 18:40:44 -07:00
Ishaan Jaffer
c9c1cbc666
fix(test_spend_management): ignore metadata.eval_information in payload comparison 2026-04-23 18:35:07 -07:00
Ishaan Jaffer
d2bfdf78e7
fix(test_gcs_pub_sub): ignore metadata.eval_information in comparison 2026-04-23 18:35:04 -07:00
Ishaan Jaffer
0348e975dc
fix(llm_as_a_judge): remove @log_guardrail_information decorator to fix duplicate guardrail_information entries
The decorator and the manual finally block both called add_standard_logging_guardrail_information_to_request_data, producing two entries per request. The decorator also misclassified HTTPException(422) blocks as guardrail_failed_to_respond (it checks for 400). The finally block correctly tracks status throughout, so removing the decorator is sufficient.
2026-04-23 17:23:13 -07:00
Ishaan Jaffer
99b0ccfad0
test(llm_as_a_judge): add unit tests for guardrail hook 2026-04-23 17:02:57 -07:00
Ishaan Jaffer
5896ec8c02
fix(llm_as_a_judge): remove dead registry dicts, fix KeyError in prompt builder, set correct status on judge failure 2026-04-23 17:02:56 -07:00
Ishaan Jaffer
7eaf1406b8
fix(LLMJudgeFields): replace @tremor/react Button with antd Button 2026-04-23 17:00:12 -07:00
Ishaan Jaffer
478558b8a9
fix(llm_as_a_judge): support Pydantic object in _get_litellm_param fallback chain 2026-04-23 17:00:11 -07:00
Ishaan Jaffer
2b4a1c4f46
fix(guardrail_endpoints): use correct PK field in rollback delete and log rollback failure 2026-04-23 17:00:10 -07:00
Ishaan Jaffer
da130d2e8e
fix(llm-as-a-judge): fix P1 code quality issues - validate weights/on_failure, guard pre_call, handle multimodal, move imports to module level, fix spurious finally logging 2026-04-23 16:25:13 -07:00
Ishaan Jaffer
af15eda10d
fix(guardrail-registry): hardcode llm_as_a_judge in initializer registry so it loads regardless of package install path 2026-04-23 16:11:25 -07:00
Ishaan Jaffer
a731d10f19
fix(guardrail-create): surface config validation errors on create instead of silently orphaning guardrail in DB 2026-04-23 15:58:51 -07:00
Ishaan Jaffer
433fdc7df6
fix(guardrail-viewer): stack lifecycle + eval details vertically to avoid badge overflow in narrow drawer 2026-04-23 15:52:52 -07:00
Ishaan Jaffer
125a8bf62c
fix(guardrails-ui): route llm_as_a_judge to criteria builder step; rename to LiteLLM LLM as a Judge; add litellm logo 2026-04-23 15:43:05 -07:00
Ishaan Jaffer
9efa01f0cd
fix(ui): wire EvalViewer into LogDetailContent to show LLM judge results on logs page 2026-04-23 15:25:04 -07:00
Ishaan Jaffer
363a4688a6
feat(ui): update EvalViewer — title 'LLM Judge Results', weighted score column, summary row 2026-04-23 15:25:04 -07:00
Ishaan Jaffer
e6ee14a249
feat(ui): wire LLM-as-a-Judge into add guardrail form 2026-04-23 15:25:04 -07:00
Ishaan Jaffer
703e9e9130
feat(ui): add LLMJudgeFields criteria builder component 2026-04-23 15:25:04 -07:00
Ishaan Jaffer
0ad29239eb
fix(a2a): filter agent-only litellm_params from acompletion kwargs; pass agent_id into body 2026-04-23 15:25:04 -07:00
Ishaan Jaffer
ffdf63df76
feat(guardrails): add self-contained llm_as_a_judge guardrail hook 2026-04-23 15:25:04 -07:00
Ishaan Jaffer
96a2b3e42f
feat(types): add EvalVerdict, StandardLoggingEvalInformation; wire eval_information into SpendLogsMetadata 2026-04-23 15:25:04 -07:00
Ishaan Jaffer
9572d4e1a0
feat(guardrails): add LLM_AS_A_JUDGE to SupportedGuardrailIntegrations 2026-04-23 15:25:04 -07:00
Michael-RZ-Berri
c81342e3c2
Merge pull request #26204 from BerriAI/litellm_budgetLimitFix
Fix bugs that bypasses per-team member budget limit
2026-04-23 14:59:07 -07:00
shin-berri
9cd1f6a599
Merge pull request #26342 from BerriAI/litellm_create_release_branch_gha
[Infra] Add standalone create-release-branch workflow
2026-04-23 14:14:00 -07:00
Mateo Wang
3950f5ea72
feat: add gpt-5.5 to model cost map (#26345)
* feat: add gpt-5.5 to model cost map

Add gpt-5.5 entry with pricing from OpenAI flagship page:
input $5/1M, cached input $0.50/1M, output $30/1M, 272K context.

* test: add gpt-5.5 coverage for model cost map and gpt-5 routing

- Add gpt-5.5 to GPT5_MODELS parametrized list so both OpenAIGPT5Config
  and AzureOpenAIGPT5Config routing tests cover the new model.
- Add test_generic_cost_per_token_gpt55 verifying the new entry's
  cost-map values ($5/$0.50/$30 per 1M) and that generic_cost_per_token
  returns the expected prompt/completion costs.
2026-04-23 14:05:22 -07:00
Michael Riad Zaky
46336b1ac3 fix linting 2026-04-23 12:08:10 -07:00
yuneng-jiang
2bfbb142b7
Merge pull request #26336 from BerriAI/litellm_yj_apr22
[IInfra] Merge dev branch
2026-04-23 12:02:47 -07:00
Yuneng Jiang
daf29d6a4a
[Infra] Add standalone create-release-branch workflow
Extracts release branch creation into a separate reusable workflow
(create-release-branch.yml) that can be triggered independently via
workflow_dispatch or called from other workflows via workflow_call.

create-release.yml now dispatches it as a dependent job after the
release publishes, keeping both workflows decoupled.
2026-04-23 12:02:25 -07:00
Yuneng Jiang
994e35135d
fix: correct image size limit enforcement and vertex_location None passthrough
token_counter.py: the previous size-limit raises were inside except Exception: pass,
so they were silently swallowed. The post-read raise was worse — img_data was already
assigned the full body before the raise, so the oversized value was used downstream.
Restructured to only assign img_data when the body is within bounds.

vertex_ai/common_utils.py and llm_passthrough_endpoints.py: the is-not-None guard
skipped validation for None, falling through to produce "https://None-aiplatform..."
Added explicit None check that raises before the regex guard.
2026-04-23 11:20:20 -07:00
yuneng-jiang
7ebe7cc99b
Merge pull request #26293 from BerriAI/litellm_img_edit_file_input_validation
[Fix] Image edit endpoints: enforce multipart-only file inputs
2026-04-23 09:42:03 -07:00
shin-berri
da3b715c36
Merge pull request #26286 from BerriAI/litellm_/unify_uv_cache
[Infra] CCI: cache, cleanup, anchors, install-path parity, Python 3.12, Ruby/Node pins
2026-04-23 09:39:06 -07:00
Sameer Kankute
d5449f5b1a
Merge pull request #26300 from BerriAI/litellm_oss_staging_04_22_2026
Litellm oss staging 04 22 2026
2026-04-23 18:53:58 +05:30
Sameer Kankute
2d1cc68e22
fix(dashscope): fail fast on image generation API errors
Prevent silent empty image responses by raising provider errors for non-200 HTTP statuses and DashScope API-level error payloads, with regression tests covering both paths.

Made-with: Cursor
2026-04-23 18:41:01 +05:30
Sameer Kankute
2e3a4bb27a
Fix black 2026-04-23 18:32:24 +05:30
Sameer Kankute
e3440baa0c
Merge pull request #25767 from vinhphamhuu-ct/main
feat: Expand VideoMetadata support to all Gemini Models.
2026-04-23 17:20:01 +05:30
Sameer Kankute
1385d46e99
FIx mypy issues 2026-04-23 17:17:06 +05:30
Yuneng Jiang
03a022436b
[Infra] CCI: run RVM install from its own checkout dir
The rvm/install script sources scripts/functions/installer using
paths relative to the caller's working directory (not $0), so
invoking /tmp/rvm/install from /home/circleci/project fails with
'No such file or directory'. Switch to (cd /tmp/rvm && ./install).
2026-04-22 21:50:53 -07:00
Yuneng Jiang
eb6a2d043c
[Infra] CCI: pin Ruby and Node.js installs in proxy_pass_through_endpoint_tests
Align the Ruby, Node.js, and npm install path with the rest of the
config. Three separate upstream installers were being invoked via
\`curl ... | bash\` or unlocked \`npm install\`:

- RVM's \`get.rvm.io/stable\` installer (mutable upstream script).
  Replace with a shallow git clone of the rvm/rvm repo at tag 1.29.12
  and verify HEAD matches the published commit SHA before running the
  local \`./install\` script. Same pattern already used for the
  helm-unittest plugin in .github/workflows/helm_unit_test.yml.
- NodeSource's \`deb.nodesource.com/setup_18.x\` piped into sudo bash.
  Replace with a direct download of the Node.js 18.20.8 linux-x64
  tarball from nodejs.org, verified against the published
  SHASUMS256.txt digest before extraction.
- \`npm install @google-cloud/vertexai @google/generative-ai\` and
  \`--save-dev jest\` resolved fresh from the npm registry on every
  run. Add \`tests/pass_through_tests/package.json\` with pinned
  direct-dep versions and commit the generated package-lock.json, then
  switch CI to \`npm ci\` (exact lockfile install, fails on drift).

Also scopes the Ruby+JS test runners to \`tests/pass_through_tests/\`
so they pick up the committed package.json rather than writing
node_modules at repo root.
2026-04-22 21:28:42 -07:00
Yuneng Jiang
a12a2190d7
[Infra] Flip remaining CI jobs to Python 3.12
Stragglers from the 2026-04-21 Python 3.12 standardization:
- .github/workflows/check_duplicate_issues.yml (was 3.11)
- .github/workflows/llm-translation-testing.yml (was 3.11)
- .github/workflows/scan_duplicate_issues.yml (was 3.13)
- .circleci proxy_build_from_pip_tests (was 3.13)

The only intentional non-3.12 CI job is installing_litellm_on_python_3_13,
which exists as an explicit "latest supported Python" smoke matrix.
2026-04-22 21:26:19 -07:00
Yuneng Jiang
547d60c642
[Infra] CCI: match Windows uv install path to Linux verification pattern
The Windows uv install step was piping a remote install.ps1 into
Invoke-Expression without any integrity check, while the Linux
install steps (install_uv command, line 89) download to a file,
verify SHA-256 against a hardcoded digest, and only then execute.
Bring the Windows path to the same pattern.

Also hardcode the kubectl v1.31.4 checksum in helm_chart_testing
instead of fetching kubectl.sha256 from the same origin as the
binary — if dl.k8s.io were ever to serve a tampered pair, a
co-hosted checksum provides no additional integrity.
2026-04-22 21:25:22 -07:00
Yuneng Jiang
44362cb167
[Infra] CCI: factor repeated filters and Python docker image to YAML anchors
The same branch filter block appeared 46 times in the workflow
declaration:

    filters:
      branches:
        only:
          - main
          - /litellm_.*/

And the same pinned Python docker image appeared 29 times in jobs:

    - image: cimg/python:3.12@sha256:9c796c...
      auth:
        username: ${DOCKERHUB_USERNAME}
        password: ${DOCKERHUB_PASSWORD}

Replace with YAML anchors declared at first use:

- `&main_branches` on using_litellm_on_windows's filters block;
  all other job entries reference it as `filters: *main_branches`.
- `&python312_image` on local_testing_part1's first docker image
  entry; all other jobs reference `- *python312_image`, including
  the multi-image jobs (auth_ui_unit_tests,
  installing_litellm_on_python_v2_migration_resolver) which keep
  their postgres sidecar entry inline afterwards.

Net result: one place to change when the image digest rolls or
the branch-filter convention changes. No behavior change — YAML
anchor resolution produces identical config at parse time.

Also adds Docker Hub auth block to upload-coverage (previously
pulled anonymously). No functional difference for a public
image, but avoids Docker Hub rate limits now that we reuse the
same entry.
2026-04-22 21:24:06 -07:00
Yuneng Jiang
bea872a034
[Infra] CCI: remove dead steps accumulated across jobs
Clean out copy-paste debug and workaround lines that serve no
purpose:

- `pwd && ls` echoes at the top of 30 "Run tests" steps (CCI
  already logs working_directory on every step).
- "Show git commit hash" in local_testing_part1/part2 and
  langfuse_logging_unit_tests (CCI shows the SHA in every job
  header).
- "Verify Docker is available" stubs in 6 machine-executor
  jobs (machine executors always have Docker).
- `sudo systemctl restart docker` in proxy_store_model_in_db_tests
  (one-off workaround; not used anywhere else).
- Duplicated Black formatting step in local_testing_part1 and
  local_testing_part2 — Black runs in the lint job, no reason to
  run it again here.
- Second back-to-back `helm test litellm --logs` invocation in
  helm_chart_testing (one call is enough).

No behavior change — these are all log-only or no-op steps.
2026-04-22 21:18:16 -07:00
Yuneng Jiang
8fbf0d5554
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/unify_uv_cache 2026-04-22 21:06:24 -07:00
Sameer Kankute
63ba912b47
Merge pull request #26160 from BerriAI/litellm_vertex_image_edit_credentials_fix
fix(image_edit): forward litellm_params to validate_environment for Vertex AI credentials
2026-04-23 08:36:01 +05:30
Zark .
fcf917df6d
Feat(dashscope): add image generation support for qwen-image-2.0 and qwen-image-2.0-pro (#25672)
* feat: add dashscope/qwen-image-2.0 and qwen-image-2.0-pro to model cost map

* feat: implement DashScope image generation transformation class

* feat: register DashScope in ProviderConfigManager for image generation

* feat: add DashScope to image generation provider routing

* feat: auto-route qwen-image /chat/completions requests to /images/generations

* test: add unit tests for DashScope image generation (22 cases)

* refactor: remove proxy-layer qwen-image auto-routing

* feat: auto-redirect image_generation models in acompletion()

* test: add acompletion auto-redirect test for image_generation models

* fix: remove unused Union import in DashScope transformation

* fix: scope acompletion redirect to dashscope and narrow exception handler

* fix: move get_str_from_messages to module-level import and forward n param to aimage_generation

* refactor: remove acompletion image_generation auto-redirect for dashscope

* test: remove acompletion auto-redirect test for dashscope image models

---------

Co-authored-by: zark.lin <zark.lin@thinkchina.com>
2026-04-22 20:03:46 -07:00
Braulio Vargas López
d26bcda52a
refactor: replace substring check with startswith in is_model_gpt_5_model (#25793)
The original check `"gpt-5-chat" not in model` already correctly
classifies all current gpt-5 variants (including gpt-5.3-chat and
gpt-5.1-chat, which do NOT contain the substring "gpt-5-chat"). This
change replaces it with an explicit `startswith("gpt-5-chat")` prefix
test on the provider-prefix-stripped model name.

The new check is functionally equivalent for all existing model names
but makes the classification boundary unambiguous and forward-safe:
future model names that might contain "gpt-5-chat" as an interior
substring won't accidentally be excluded from the GPT-5 reasoning path.

Also moves the new regression test from tests/ root to
tests/test_litellm/llms/openai/ so it is included in `make test-unit`.
2026-04-22 19:55:00 -07:00
Rick
c26e304abc
fix(ui): stale filters applied after sort/page/time change on Request Logs (#25789)
The useEffect that re-fetches logs on sort/page/time changes:

  useEffect(() => {
    if (hasBackendFilters && accessToken) {
      performSearch(filters, currentPage);
    }
  }, [sortBy, sortOrder, currentPage, startTime, endTime, isCustomDate]);

intentionally omits `filters` and `hasBackendFilters` from its dep array
to avoid double-fetches when a filter is applied.  The side-effect is a
stale-closure bug: the effect captures `filters` and `hasBackendFilters`
from the render where its deps last changed, not from the render where
the user selected, e.g., a Key Alias.

Reproduce: set Key Alias → results appear correctly → change page or
sort → the effect fires with the OLD `filters` snapshot (no key_alias)
→ API request is sent without the filter → table shows unfiltered data.

Fix: store the latest `filters` and `hasBackendFilters` in refs that are
kept in sync on every render.  The sort/page/time effect reads from the
refs instead of the closure so it always uses the current filter state
without altering the dep array.

Co-authored-by: Bytechoreographer <Bytechoreographer@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4 (1M context) <noreply@anthropic.com>
2026-04-22 19:41:34 -07:00
Rick
4b2fd870ca
fix(ui): Fetch button ignores active filters on Request Logs page (#25788)
When backend filters (e.g. Key Alias) are active on the Request Logs
page, the manual Fetch button called logs.refetch() which re-runs the
main TanStack Query.  That query does not carry backend-only filter
params such as key_alias, so the button had two problems:

1. It fired a redundant API request without the active filters.
2. It did not refresh the filtered result set — backendFilteredLogs
   stayed frozen at the last debounce-triggered fetch.

Fix: expose refetchWithFilters() from useLogFilterLogic and route the
Fetch button through it when hasBackendFilters is true.  This cancels
any in-flight debounce and calls performSearch with the current filter
state, keeping all active filters intact.

Co-authored-by: Bytechoreographer <Bytechoreographer@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4 (1M context) <noreply@anthropic.com>
2026-04-22 19:39:24 -07:00
BillionToken
947931858e
fix(anthropic): handle tool_choice type 'none' in messages API (#24457)
* fix(anthropic): handle tool_choice type 'none' in messages API

* test(anthropic): add regression test for tool_choice type 'none'

---------

Co-authored-by: BillionClaw <267901332+BillionClaw@users.noreply.github.com>
Co-authored-by: Krrish Dholakia <krrish+github@berri.ai>
2026-04-22 19:36:18 -07:00
Elias
bd145d18e1
fix(ovhcloud): Fix tool calling not working (#25948)
* fix(ovhcloud): fix tool calling

* fix import order
2026-04-22 19:33:58 -07:00