Commit graph

27064 commits

Author SHA1 Message Date
Fabrício Ceschin
387982fcbf
Fix langfuse input tokens logic for cached tokens (#16203)
* Update langfuse.py

Fixing issue with input_tokens and cache_read_tokens

* Clarify input token calculation in langfuse.py

Add comment to clarify input token calculation based on Langfuse documentation.
2025-11-04 19:20:35 -08:00
Alexsander Hamir
727fe504bb
fix(redis): handle float redis_version from AWS ElastiCache Valkey (#16207)
* fix(redis): handle float redis_version from AWS ElastiCache Valkey

AWS ElastiCache Valkey returns redis_version as a float (7.0) instead
of a string ('7.0.0'), causing AttributeError: 'float' object has no
attribute 'split' in async_lpop when parsing version for LPOP count.

Changes:
- Extract version parsing into _parse_redis_major_version() helper
- Add DEFAULT_REDIS_MAJOR_VERSION constant (replaces magic number)
- Support multiple version formats: string, float, int, malformed
- Add comprehensive test coverage for all version format edge cases

Fixes: 'LiteLLM Redis Cache LPOP: - Got exception from REDIS' error
during db_spend_update_job cronjobs

* refactor: move DEFAULT_REDIS_MAJOR_VERSION to constants.py
2025-11-04 19:20:00 -08:00
Cesar Garcia
44a928b631
fix(openai): Remove automatic summary from reasoning_effort transformation (#16210)
* Fix: Remove automatic summary field from reasoning_effort transformation

Problem:
The _map_reasoning_effort() function was automatically adding
reasoning.summary field when users specified reasoning_effort parameter,
causing 400 errors for users with unverified OpenAI organizations.

Root Cause:
According to OpenAI's official documentation, the summary field is opt-in
and requires organization verification:

"Reasoning summary output [...] will not be included unless you explicitly
opt in to including reasoning summaries."

"Before using summarizers with our latest reasoning models, you may need
to complete organization verification"

Source: https://platform.openai.com/docs/guides/reasoning#reasoning-summaries

Solution:
Remove the automatic inclusion of summary field from all reasoning_effort
levels (high, medium, low, minimal). Users who want reasoning summaries
can explicitly pass reasoning={"effort": "high", "summary": "auto"} in
their requests.

Impact:
- Fixes #16032
- Works for all organizations (verified and unverified)
- Maintains backward compatibility for users passing reasoning object directly
- Follows OpenAI's recommended opt-in approach

Testing:
- All existing tests pass (4/4 tests in transformation suite)
- Manual verification confirms only effort field is included

* test: Fix MockResponse missing headers attribute in test_openai_responses_api

The MockResponse class was missing the 'headers' attribute which caused
APIConnectionError when processing the mock response. Added headers={}
to fix the test.

* feat: Add dict support to reasoning_effort parameter

Allow users to pass reasoning_effort as either:
- String: reasoning_effort="high" (no summary, safe default)
- Dict: reasoning_effort={"effort": "high", "summary": "detailed"} (opt-in)

This preserves backward compatibility while giving users flexibility
to explicitly opt-in to the summary field when needed (for verified
OpenAI organizations).

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2025-11-04 19:05:17 -08:00
Bowen Liang
4e12e3f90d
fix typo of orginal (#16255) 2025-11-04 18:55:44 -08:00
Alexsander Hamir
b4def8899f
add: shared_session support to responses API (#16260)
This change enables HTTP client session reuse by:
- Adding shared_session parameter to all responses API methods (responses, delete_responses, get_responses, list_input_items, cancel_responses)
- Passing shared_session to get_async_httpx_client() for connection pooling
- Adding debug logging to track shared session usage

This helps reduce memory overhead by reusing HTTP connections instead of creating new clients for each request, which is particularly important for high-throughput proxy scenarios.
2025-11-04 18:53:54 -08:00
Alan Ponnachan
b564ed81f5
add reasoning effort and test (#16261) 2025-11-04 18:53:02 -08:00
Ishaan Jaffer
4a83ae0695 test_aresponses_service_tier_and_safety_identifier 2025-11-04 18:05:30 -08:00
Ishaan Jaffer
ede8944e49 fix _handle_callback_failure 2025-11-04 18:00:12 -08:00
Ishaan Jaffer
5b658ce1eb ui fix 2025-11-04 17:58:51 -08:00
Ishaan Jaff
60f3a3b0ad
[Feat] add serxng search API provider (#16259)
* TestFirecrawlSearch

* add SearchProviders

* add to get_provider_search_config

* add FirecrawlSearchConfig

* add FirecrawlSearchRequest

* add firecrawl API docs

* add pricing firecrawl/search

* add new search APIs

* add SearXNGSearchConfig

* add searxng/search

* add serxng params

* TestSearXNGSearch

* docs serxng

* docs fix

* docs fix

* docs serxng
2025-11-04 17:56:07 -08:00
Ishaan Jaff
af78a93ecf
[Feat] /search API - add firecrawl search API support (#16257)
* TestFirecrawlSearch

* add SearchProviders

* add to get_provider_search_config

* add FirecrawlSearchConfig

* add FirecrawlSearchRequest

* add firecrawl API docs

* add pricing firecrawl/search

* add new search APIs
2025-11-04 17:52:12 -08:00
Krrish Dholakia
a3810a5f3a docs(openai_passthrough.md): document how to make openai passthrough route work 2025-11-04 17:12:33 -08:00
Ishaan Jaff
c1ac6e25aa
[Feat] Add Bedrock Agentcore as a provider on LiteLLM Python SDK and LiteLLM AI Gateway (#16252)
* add agentcore in get_bedrock_route

* add AmazonAgentCoreConfig

* fix get_runtime_endpoint

* init AmazonAgentCoreConfig

* add get_bedrock_chat_config

* get_bedrock_chat_config

* add AmazonAgentCoreConfig

* fix get_complete_url

* refactor transform response

* test agentcore

* test_bedrock_agentcore_with_streaming

* fix _parse_json_response

* fix _calculate_usage

* test_bedrock_agentcore_basic

* add AgentCoreSSEStreamIterator

* add native streaming for agentcore

* test_bedrock_agentcore_with_streaming

* test_bedrock_agentcore_basic

* add agentcore

* _calculate_usage

* fix linting
2025-11-04 16:35:12 -08:00
Deepanshu Lulla
812ea03d28
Add tags and descriptions support to aws secrets manager (#16224)
* Add tags and descriptions support to aws secrets manager

* add tags

---------

Co-authored-by: deepanshu <deepanshu.lulla@hq.bill.com>
2025-11-04 16:11:51 -08:00
yuneng-jiang
2b0aea87c0
[Feature] UI - Initial changes for supporting prompts to multiple models (#16223)
* Initial changes for supporting prompts to multiple models

* Additional tests to prevent regressions

* Resolve merge conflicts

* Fixing E2E build issues
2025-11-04 16:02:19 -08:00
Cesar Garcia
78ed5126a5
fix: Fix Responses API streaming tests usage field names and cost (#16236)
This commit fixes two bugs in Responses API streaming tests:

1. **Usage field naming bug**: Tests were using `input_tokens` and
   `output_tokens` but the Usage object uses `prompt_tokens` and
   `completion_tokens`.

2. **Missing cost in streaming usage**: When `include_cost_in_streaming_usage`
   was enabled, the cost was calculated and added to ResponseAPIUsage, but was
   lost during the transformation to the Usage object.

Changes:
- Updated test assertions to use correct field names (prompt_tokens, completion_tokens)
- Added cost preservation logic in FakeStreamerResponsesAPIIterator
- Modified _transform_response_api_usage_to_chat_usage() to preserve cost attribute

All streaming tests now pass successfully.
2025-11-04 15:57:59 -08:00
Alexsander Hamir
7878ebd2a2
fix(proxy): handle None values in daily spend sort key (#16245)
Fixes TypeError when sorting daily spend transactions with missing entity IDs (e.g., null tags).
Changed x[1].get(entity_id_field) to x[1].get(entity_id_field) or "" to provide a sortable
default value when the field is None, preventing comparison errors between NoneType and str.
2025-11-04 15:56:48 -08:00
yuneng-jiang
2ed95e216c
[Feature] UI - Tag Usage Top Model Table View and Label Fix (#16249)
* Tag Usage Top Model Table View and Label Fix

* Removed unused state
2025-11-04 14:24:43 -08:00
yuneng-jiang
4d756a62d9
Prevent trailing slash in sso proxy base url input (#16244) 2025-11-04 14:22:50 -08:00
yuneng-jiang
615f76de88
[Feature] UI - Litellm test key audio (#16251)
* Add audio/transcriptions and audio/speech to Test key UI

* Removed unused import in test

* Fixed voice parameter not being passed in
2025-11-04 14:21:26 -08:00
Krish Dholakia
726452dfe8
Revert "add: comparison with portkey (#16145)" (#16247)
This reverts commit 844ace1283.
2025-11-04 10:37:56 -08:00
Alexsander Hamir
844ace1283
add: comparison with portkey (#16145) 2025-11-04 10:37:17 -08:00
YutaSaito
8e27b6c0b4
[MCP] configure static mcp header (#16179)
* feat: configure extra mcp headers in ui

* doc: static header

* build: add new migration file

* chore: add missing image file

* fix: test
2025-11-03 21:06:36 -08:00
Alan Ponnachan
45cf6fc289
feat: support dashscope tiered pricing
* add helper functions

* update generic_cost_per_token function

* add test

* formatting

* add examples in docstring for _calculate_tiered_cost

* Restore files to upstream/main version

* dashscope specific calculation

* improve for different costs

* remove _calculate_flat_cost function
2025-11-03 20:27:25 -08:00
pablobgar
5ddc5410bf
fix: cumulative index (#16194) 2025-11-03 20:25:05 -08:00
Sameer Kankute
8a904a5481
Add gemini live audio model cost in model map (#16183)
* Add gemini live audio model cost in model map

* add gemini models
2025-11-03 19:01:00 -08:00
Anthony Ivan
e5e9523958
init commit (#16200) 2025-11-03 18:58:03 -08:00
Niv Goldenberg
232d1558dd
fix(anthropic-adapter): properly translate Anthropic image format to OpenAI (#16202)
* fix(anthropic-adapter): properly translate Anthropic image format to OpenAI

Fixed bug where images were stripped during Anthropic Messages API to Azure
OpenAI translation. Image source data was being stringified instead of having
fields properly extracted.

- Added _translate_anthropic_image_to_openai() helper method
- Support both base64 and URL image formats per Anthropic API spec
- Refactored user message and tool result image handling

* test(anthropic-adapter): add comprehensive image translation tests

Add 5 unit tests covering image translation from Anthropic to OpenAI format:
- User messages with base64 images
- User messages with URL images
- Tool results with base64 images
- Tool results with URL images
- Mixed content with multiple images
2025-11-03 18:53:44 -08:00
Sameer Kankute
bb86c94df4
Add Prometheus metric to track callback logging failures in S3 (#16209)
* Add v1 cut of container api

* fix lint errors

* Add proxy support to container apis & logging support (#16049)

* Add proxy support to container apis

* Add logging support

* Add cost tracking support for containers and documentation

* Add new constant documentation

* Add container cost in model map

* fix failing azure tests

* Update tests based on model map changes

* fix model map tests

* fix model map tests

* Container modeshould be container

* Container tests fix

* Merge branch 'main' into litellm_sameer_oct_staging_2

* Add Prometheus metric to track callback logging failures in S3 (#16102)

* Add proxy support to container apis

* Add logging support

* prometheus metric  measures how often s3_v2 is failing

* remove not needed files

* remove not needed files

* remove not needed files

* fix mypy errors

* Use logging_callback_manager to get all the callbacks

---------

Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
2025-11-03 18:46:52 -08:00
Krrish Dholakia
db1b38381a fix(llm_passthrough_endpoints.py): merge query params correctly for gemini 2025-11-03 18:30:37 -08:00
Ishaan Jaffer
4365fc57ee fix SpendLogsMetadata 2025-11-03 18:08:35 -08:00
Ishaan Jaffer
f09736ceb9 fix mypy lint 2025-11-03 18:02:19 -08:00
Ishaan Jaffer
9b4b846af9 refactor 2025-11-03 17:31:55 -08:00
Ishaan Jaffer
ed8235ea07 fix typing 2025-11-03 17:25:49 -08:00
Ishaan Jaff
57295cedef
[Feat] Add Azure AI Doc Intelligence OCR (#16219)
* TestAzureDocumentIntelligenceOCR

* add AZURE_DOCUMENT_INTELLIGENCE_API_VERSION

* add AzureDocumentIntelligenceOCRConfig

* add async_transform_ocr_response

* use async transform

* add AzureDocumentIntelligenceOCRConfig

* add AzureDocumentIntelligenceOCRConfig

* add AzureDocumentIntelligenceOCRConfig

* add get_azure_ai_ocr_config

* add azure_ai/doc-intelligence

* add azure_ai/doc-intelligence

* docs fix

* docs fix

* add azure doc intel

* fix lint error
2025-11-03 17:22:19 -08:00
Alexsander Hamir
a73e890d8f
fix: broken link on model_management.md (#16217) 2025-11-03 17:00:03 -08:00
Ishaan Jaff
71c61c274f
[Feat] /ocr - Add VertexAI OCR provider support + cost tracking (#16216)
* add VertexAIOCRConfig

* __all__ = ["VertexAIOCRConfig"]
add

* add get_provider_ocr_config

* use GenericLiteLLMParams for litellm params

* fix _async_prepare_ocr_request

* fix _prepare_ocr_request

* fix get_complete_url

* fix validate_environment

* add safe_get_vertex_ai_project

* add VertexAIOCRConfig

* fix get_complete_url

* add TestVertexAIOCR

* add mistral-ocr-2505 cost

* add OCR to provider info

* docs vertex ai ocr

* fix _handle_rate_limits

* Potential fix for code scanning alert no. 3632: Clear-text logging of sensitive information

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2025-11-03 15:56:49 -08:00
Ishaan Jaff
40e96657ce
revert noma apply guard (#16214) 2025-11-03 14:57:40 -08:00
Ishaan Jaff
0737cc7c13
[Feat] s3 logger, add support for ssl_verify when using minio logger (#16211)
* fixes s3_v2 verify

* test_s3_verify_false_async_client

* fix

* ruff fixes
2025-11-03 13:56:00 -08:00
Sameer Kankute
df6e084984
Fix image_config.aspect_ratio not working for gemini-2.5-flash-image (#15999)
* fix image edit method

* fix mypy error
2025-11-03 08:48:36 -08:00
Sameer Kankute
ad6a0f4d44
Update perplexity cost tracking (#15743)
* Update perplexity cost tracking

* fix lint errors

* fix code

* fix tests in perplexity

* fix test realted to api call

* fix exception test
2025-11-03 08:45:34 -08:00
Sameer Kankute
396ab80f56
Fix index field not populated in streaming mode with n>1 and tool calls (#15962)
* fix index tool calling in streaming

* moved test to llm translation

---------

Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
2025-11-02 09:52:31 -08:00
Krish Dholakia
07d2a27f14
Milvus - Passthrough API support - adds create + read vector store support via passthrough API's (#16170)
* feat(llm_passthrough_endpoints.py): support milvus passthrough api

* fix(llm_passthrough_endpoints.py): move streaming request value to the top of the function

* docs: document new milvus vector store passthrough flow
2025-11-02 09:47:58 -08:00
YutaSaito
6ed76ff809
feat: change guardrail_information to list type (#16127)
* feat: change guardrail_information to list type to support displaying multiple guardrails

* fix: add missing commit and revert auto-format changes in utils.py

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2025-11-02 09:47:49 -08:00
Krish Dholakia
74ae7aed44
build: Squashed commit of the following: (#16176)
commit bb0b050fb0
Author: Krrish Dholakia <krrishdholakia@gmail.com>
Date:   Sat Nov 1 20:00:01 2025 -0700

    test: update tests

commit b2da4bdac2
Author: Krrish Dholakia <krrishdholakia@gmail.com>
Date:   Wed Oct 22 14:58:01 2025 -0700

    fix(langfuse_otel_attributes.py): log tools and other optional params

commit 75bee1f274
Author: Krrish Dholakia <krrishdholakia@gmail.com>
Date:   Wed Oct 22 14:42:05 2025 -0700

    feat(langfuse_otel/): working request/response logging on spans

    Closes https://github.com/BerriAI/litellm/issues/13764

commit a3e4fa5b81
Author: Krrish Dholakia <krrishdholakia@gmail.com>
Date:   Wed Oct 22 14:20:39 2025 -0700

    fix: initial commit fixing langfuse request/response logging with OTEL

commit 09fc9deac8
Author: Krrish Dholakia <krrishdholakia@gmail.com>
Date:   Wed Oct 22 13:33:52 2025 -0700

    fix(litellm_logging.py): for responses api - return a unified usage object for logging

    ensures logging integrations all pull the right usage information
2025-11-02 09:46:40 -08:00
Geoffray Viossat
3922bb6ed5
fix: return the diarized transcript when it's required in the request (#16133) 2025-11-02 09:45:18 -08:00
Katsuhiro Muto
99775fa0f8
Support responses API streaming in langfuse otel (#16153)
* streaming support in langfuse otel

* Added testing for Langfuse Otel tracing in the response API

---------

Co-authored-by: eycjur <eycjur@example.com>
2025-11-02 09:36:34 -08:00
Krish Dholakia
3f40613c56
fix(ui_sso.py): support dot notation on ui sso (#16135) 2025-11-02 09:35:52 -08:00
Deepanshu Lulla
20b95e9a80
strip base64 in s3 (#16157)
* strip base64

* strip base64

* s3 use key prefix

* s3 use key prefix

* strip base64 doc

---------

Co-authored-by: deepanshu <deepanshu.lulla@hq.bill.com>
2025-11-02 09:06:53 -08:00
yuneng-jiang
e336434b27
[Feature] UI - Guardrail Info Page Show PII Config (#16164)
* Guardrail info page fix

* Make the Configurations more readable and render in a table
2025-11-02 09:05:00 -08:00