Commit graph

36888 commits

Author SHA1 Message Date
Albert Sebastian
e3cda5f919 fix(fal_ai): prevent double fal-ai/ prefix in generic model URL construction
When model is configured as fal_ai/fal-ai/model-name, the provider prefix
stripping leaves fal-ai/ intact, causing get_complete_url() to produce
fal-ai/fal-ai/model-name. Use removeprefix() to strip it before re-adding.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:56:09 +05:30
Albert Sebastian
d90af8c8bc fix(ui): rebuild dashboard to include image preview for fal_ai logs
The pre-built UI static files were stale — source changes for image
generation response rendering (OutputCard, prettyMessagesUtils) were
not compiled into _experimental/out/.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:43:23 +05:30
Albert Sebastian
926e2e641c feat: add resolution-based and per-megapixel pricing for FAL models
Enhance the FAL AI cost calculator to support three pricing modes:
- PER_MEGAPIXEL: cost scales with image dimensions (flux-2/klein models)
- PER_IMAGE_RESOLUTION: different rates for 2K vs 4K (nano-banana, gemini, seedream)
- PER_CALL: flat per-image rate (default, backward compatible)

Add pricing_basis, output_cost_per_megapixel, and output_cost_per_image_by_resolution
fields to ModelInfo. Update pricing JSON with correct USD values (1 credit = $0.01).
Set response size from request params in FAL transformation for cost calculation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:42:14 +05:30
Albert Sebastian
ea84e98c78 feat: add fal_ai model pricing and image preview in log viewer
Add pricing for 23 fal_ai image generation models (nano-banana, flux, gemini,
seedream, kling, iclight). Display generated images inline in the log details
drawer instead of raw JSON for image generation responses.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:42:14 +05:30
Albert Sebastian
5f1d3c1fa8 fix(fal_ai): revert get_supported_openai_params to standard params only
FAL-specific params (image_urls, loras, etc.) must NOT be listed in
get_supported_openai_params(). LiteLLM's _get_non_default_params() only
processes keys matching default_params (n, quality, size, style, user).
Params listed in get_supported_openai_params but not in default_params
are excluded from BOTH the standard path AND the provider-specific
safety net (add_provider_specific_params_to_optional_params), causing
them to be silently dropped.

By keeping only standard OpenAI params, FAL-specific params flow through
the provider-specific safety net and reach the request body correctly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:42:14 +05:30
Albert Sebastian
3d0b82e6db fix(fal_ai): flatten extra_body in image generation request transform
LiteLLM's image_generation() path doesn't extract extra_body from kwargs
(unlike image_edit), so params like image_urls, loras, enable_safety_checker
arrive as a nested dict in optional_params. Flatten extra_body in both
map_openai_params() and transform_image_generation_request() so FAL-specific
params appear at the top level of the request body.

Adds comprehensive test coverage (35 tests):
- URL construction for all model types
- Param passthrough (loras, guidance_scale, etc.)
- extra_body flattening (the bug scenario)
- End-to-end request validation with mocked HTTP

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:42:14 +05:30
Albert Sebastian
407aba879c fix(fal_ai): fix URL construction and param passthrough for generic FAL models
Generic FAL models (nano-banana-2, flux-2/klein, etc.) were getting
incorrect URLs (just https://fal.run without model path) and having
FAL-specific params (loras, image_urls, guidance_scale, etc.) silently
dropped by the supported params filter. This broke cost tracking since
all generic FAL requests failed through the proxy.

Also adds FAL image edit config and k8s dev ConfigMap.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:42:14 +05:30
Albert Sebastian
acd64b9d54 ci: harden build-push workflow with pinned actions and permissions
Pin all GitHub Actions to commit SHAs and add top-level permissions
block to address security scanner findings (zizmor + CodeQL).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:42:14 +05:30
Albert Sebastian
fd05892dd6 ci: add GitHub Actions workflow to build and push to Azure ACR
Builds the LiteLLM Docker image on push to main and fix/* branches,
then pushes to Azure Container Registry with commit SHA and latest tags.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 09:42:14 +05:30
Albert Sebastian
b60e30a58b feat(vertex_ai): support image_size (2K/4K) for Gemini image generation
Fixes #24621

The Gemini API supports `imageConfig.imageSize` to control output
resolution (e.g., "2K", "4K"), but LiteLLM had no way to pass this
through. The `extra_body` approach doesn't work because
`generationConfig` is rebuilt from scratch in the transformation layer.

Changes:
- Vertex AI Gemini image edit: add `imageSize` to `SUPPORTED_PARAMS`
  and include it in `generationConfig.image_config.image_size`
- Google AI Studio image gen: same support for Gemini models
- Both: accept `size` param as either OpenAI format ("1024x1024" ->
  aspect_ratio) or Gemini format ("2K" -> image_size)
- Both: map OpenAI `quality="hd"` to `imageSize="2K"` as a convenient
  alternative
- Add tests for imageSize, combined aspect_ratio+imageSize, size="2K",
  and quality="hd" mappings
2026-04-14 09:42:13 +05:30
Sameer Kankute
5e80e075c7
Merge pull request #25397 from BerriAI/litellm_oss_staging_04_08_2026
Litellm oss staging 04 08 2026
2026-04-13 09:12:40 +05:30
Sameer Kankute
fa605d85c0
Merge pull request #25616 from BerriAI/main
merge main
2026-04-13 08:43:43 +05:30
yuneng-jiang
5544803b35
Merge pull request #25406 from BerriAI/litellm_regen_key_modal_antd
Some checks are pending
CodeQL / Analyze (actions) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
CodSpeed Benchmarks / benchmarks (push) Waiting to run
Helm unit test / unit-test (push) Waiting to run
Read Version from pyproject.toml / read-version (push) Waiting to run
Scorecard supply-chain security / Scorecard analysis (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Security / security (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
[Refactor] UI - Virtual Keys: migrate regenerate key modal to AntD
2026-04-11 21:27:09 -07:00
Yuneng Jiang
1857be43a7
Merge remote-tracking branch 'origin/main' into litellm_regen_key_modal_antd 2026-04-11 20:45:17 -07:00
ishaan-berri
fdd7500904
blog: add back arrow to blog post pages (#25587)
* blog: add back arrow to post pages

* blog: style back arrow — fixed top-left below navbar
2026-04-11 19:15:45 -07:00
ishaan-berri
1edf41c26f
Merge pull request #25585 from BerriAI/litellm_dev_04_11_2026_p1
Litellm dev 04 11 2026 p1
2026-04-11 18:46:57 -07:00
Krrish Dholakia
973986aac2 docs: readme tweak 2026-04-11 18:34:23 -07:00
ishaan-berri
12c1467228
Merge pull request #25583 from BerriAI/blog/ramp-style-restyle-with-redis-post
blog: Ramp-style engineering blog restyle + Redis circuit breaker post
2026-04-11 18:31:34 -07:00
Ishaan Jaffer
35f4b47ff8
apply content guidelines: scale/resilience narrative, FAQ, Key Takeaways, Conclusion CTA 2026-04-11 18:12:32 -07:00
ishaan-berri
f74d626253
Merge pull request #25580 from BerriAI/blog/ramp-style-restyle-with-redis-post
blog: restyle docs.litellm.ai/blog to engineering blog aesthetic
2026-04-11 18:10:57 -07:00
Ishaan Jaffer
14eed24471
add redis circuit breaker blog post with React diagrams 2026-04-11 18:02:59 -07:00
Ishaan Jaffer
8e616ecdf4
add BlogPostPage swizzle: hide sidebar, add hiring CTA on every post 2026-04-11 18:02:56 -07:00
Ishaan Jaffer
dac44fb443
blog list styles: clean typography, marquee animation, hero layout 2026-04-11 18:02:52 -07:00
Ishaan Jaffer
85cb7db8b9
blog list page: Ramp-style flat list with hero, provider marquee, hiring CTA 2026-04-11 18:02:48 -07:00
Ishaan Jaffer
05d516482f
restyle blog list page to match engineering blog aesthetic 2026-04-11 18:02:44 -07:00
Yuneng Jiang
7c6bd98fdf
fix: address PR review comments on RegenerateKeyModal
- Remove redundant useEffect cleanup that duplicated handleClose logic
- Remove unnecessary currentAccessToken state, use accessToken from hook directly
2026-04-11 18:01:02 -07:00
Krrish Dholakia
e08e3bf748 docs: clarify how to get benchmarking script 2026-04-11 17:31:03 -07:00
Krrish Dholakia
12bca649fc docs: refactor benchmarking docs to be clearer 2026-04-11 17:30:09 -07:00
yuneng-jiang
2c786ca2e6
Merge pull request #25578 from BerriAI/yj_apr11_bump
Some checks are pending
CodeQL / Analyze (actions) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
CodSpeed Benchmarks / benchmarks (push) Waiting to run
Helm unit test / unit-test (push) Waiting to run
Read Version from pyproject.toml / read-version (push) Waiting to run
Scorecard supply-chain security / Scorecard analysis (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Security / security (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
bump: version 1.83.6 → 1.83.7
2026-04-11 17:09:31 -07:00
Yuneng Jiang
e162c6d502
bump: version 1.83.6 → 1.83.7 2026-04-11 17:00:33 -07:00
yuneng-jiang
d7823e3a09
Merge pull request #25577 from BerriAI/yj_apr11_build_3
[Infra] Rebuild UI
2026-04-11 16:04:12 -07:00
Yuneng Jiang
a89ad3aaff
chore: update Next.js build artifacts (2026-04-11 22:57 UTC, node v22.16.0) 2026-04-11 15:57:01 -07:00
yuneng-jiang
99cb3af1a6
Merge pull request #25562 from BerriAI/litellm_internal_staging_04_11_2026
Litellm internal staging 04 11 2026
2026-04-11 15:55:29 -07:00
Yuneng Jiang
c40e459447
fix linting 2026-04-11 15:44:15 -07:00
Yuneng Jiang
909247785e
Merge remote-tracking branch 'origin' into litellm_internal_staging_04_11_2026 2026-04-11 15:41:03 -07:00
yuneng-jiang
9a43e32d6e
Merge pull request #25464 from BerriAI/litellm_passthrough_contenttype
fix(proxy): pass-through multipart uploads and Bedrock JSON body
2026-04-11 15:40:16 -07:00
shivam
3742b0cee1
refactor: define pass-through custom body state key in types module
Avoid module-level cyclic import between llm_passthrough_endpoints and
pass_through_endpoints; CodeQL and partial init order no longer risk
undefined LITELLM_PASS_THROUGH_CUSTOM_BODY_STATE_KEY.

Made-with: Cursor
2026-04-11 15:26:44 -07:00
shivam
574eb207c8
fix: import LITELLM_PASS_THROUGH_CUSTOM_BODY_STATE_KEY for bedrock passthrough
Fixes NameError when bedrock_proxy_route sets custom body on request.state.
Remove unused lazy-loader helper.

Made-with: Cursor
2026-04-11 15:08:10 -07:00
Shivam Rawat
627d70affd
Potential fix for pull request finding 'CodeQL / Module-level cyclic import'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-04-11 14:46:02 -07:00
yuneng-jiang
21e4071f25
Merge pull request #25573 from BerriAI/yj_apr11_build_2
[Infra] Rebuild UI
2026-04-11 14:25:44 -07:00
shivam
20b16b377c
chore(proxy): match CRLF line endings in pass_through_endpoints
Restores Windows-style line endings to match main/origin main for this
file, removing the full-file noise diff from an accidental LF-only
normalization.

Made-with: Cursor
2026-04-11 14:24:26 -07:00
Yuneng Jiang
b42a8750df
chore: update Next.js build artifacts (2026-04-11 21:00 UTC, node v22.16.0) 2026-04-11 14:00:21 -07:00
yuneng-jiang
56e82451df
Merge pull request #25458 from BerriAI/litellm_team_members_logs
Team member permission /spend/logs for team-wide spend logs (UI + RBAC)
2026-04-11 13:55:00 -07:00
yuneng-jiang
dbcdb1565c
Merge pull request #25241 from BerriAI/es_allow_applyguardrails_for_iam
added applyguardrail to inline iam
2026-04-11 13:51:45 -07:00
yuneng-jiang
0169f140bf
Merge pull request #25571 from BerriAI/yj_apr11_build
[Infra] Build UI for release
2026-04-11 13:31:23 -07:00
ishaan-berri
da75012a50
Merge pull request #25569 from BerriAI/litellm_harish_april11
Litellm harish april11
2026-04-11 13:26:18 -07:00
Yuneng Jiang
18caee3ef4
chore: update Next.js build artifacts (2026-04-11 20:19 UTC, node v22.16.0) 2026-04-11 13:19:41 -07:00
harish-berri
ec0cd5c17d
Update streaming.py. Provide Type Annotation for empty dict 2026-04-11 13:04:25 -07:00
yuneng-jiang
150c37c47d
Merge pull request #25568 from BerriAI/litellm_yj_apr_10_2026
[Infra] Merge dev with main
2026-04-11 13:03:52 -07:00
Yuneng Jiang
218daca867
[Fix] Address Greptile review: POST /organization/info auth bypass, inline imports, team access denial tests
- Add _verify_org_access to deprecated POST /organization/info endpoint
- Move get_user_object to module-level import in organization_endpoints.py
- Add tests for _verify_team_access 403 denial path
2026-04-11 12:40:55 -07:00