* feat(langfuse_otel): add Langfuse OpenTelemetry integration for observability
- Introduced a new integration for Langfuse OpenTelemetry, allowing users to send LiteLLM traces and observability data.
- Updated sidebars to include documentation for the new integration.
- Added example usage and configuration details in the documentation.
- Implemented necessary classes and methods to handle OpenTelemetry attributes and configuration.
- Included tests to validate the integration functionality and environment variable handling.
Still WIP
* Remove example script for Langfuse OpenTelemetry integration with LiteLLM
* feat(enterprise/): fix remaining users check on license
* fix(usage_indicator.tsx): if no max user set, don't render remaining user info card
only for users with user limits on their license
* fix(leftnav.tsx): only show remaining users to admin
* feat(columns.tsx): don't allow sorting on model access groups
it's a list[str]
* feat(model_dashboard.tsx): add model access group filters
* docs(index.md): add stable pip package
* fix(anthropic/chat/transformation.py): add 'none' tool choice mapping
Allows disabling anthropic tool calling
Maintain parity
* fix(transformation.py): if tool_choice="none" ignore 'disable_parallel_Tool_use'
unsupported param from anthropic - makes sense as the 'none' implies no tool calls are being made
* fix(anthropic/chat/transformation.py): append prefix to start of assistant response, if set
ensures assistant response contains complete response
* fix(anthropic/chat/transformation.py): add flag to allow user to opt out of enabling prefix in prompt
* fix(anthropic/chat/transformation.py): working e2e support for prefix prompt in assistant response
* feat(networking.tsx): always include model access groups on UI
show admin created access groups when giving key/user/team model permissions
* feat(add_model_tab.tsx): initial ui component for adding to an existing model access group
allows user to add model to an access group (simplify giving users/keys/teams model access)
* feat(proxy_server.py): add 'only_model_access_groups' flag support to `/v1/models`
simplifies listing available access groups on UI
* test: add e2e test for new only_model_access_groups param
* feat(add_model_tab.tsx): allow adding+viewing model access groups on models tab
make feature functional on UI
* feat(view_users.tsx): route edit user to user info page
more detailed user edit
* feat(columns.tsx): route edit user to user info page
more detailed user edit
* fix(columns.tsx): fix linting error
* build(ui/): fix linting errors
* feat(anthropic/passthrough): pass dynamic api key/api base params to litellm.completion
allows calls to work with config.yaml
* fix(responses_api/transformation): fix passing dynamic params to responses api from .completion()
Allows responses api to work with config.yaml
* fix(langfuse.py): fix responses api usage logging to langfuse
* refactor(litellm_logging.py): add more generic solution for responses api usage logging
ensures it works across all logging integrations
* fix(litellm_logging.py): patch for anthropic messages not returning a pydantic object
it should ideally return a pydantic object, which would simplify checks and reduce errors
* fix(handler.py): correctly bubble up empty choices errors to litellm.completion
causes downstream errors as it is expected there is at least one choice set
* feat(litellm_logging.py): prevent double logging litellm responses
ensures accurate spend tracking for calls when bridges are used
* fix(litellm_logging.py): ensure logging is consistently enforced across all call types
* fix: patch - set calltype before entering bridge api
ensures logging object is applying the correct logic on the event hooks
* fix(types/router.py): loosen type hint for mock response
* change space_key header to space_id for Arize (#11595)
* feat(schema): add additional indexes to LiteLLM_SpendLogs for improved query performance (#11675)
* Revert "feat(schema): add additional indexes to LiteLLM_SpendLogs for improve…" (#11683)
This reverts commit 2a7f113fde.
* [Feat] Use dedicated Rest endpoints for list, calling MCP tools (#11684)
* fix: (fix) use specific rest endpoints for MCP
* ui - use rest mcp endpoints
* fix imports
* docs DISABLE_AIOHTTP_TRUST_ENV
* docs(caching.md): remove batch redis get recommendation - old code path, no longer necessary
* fix(vertex_and_google_ai_studio_gemini.py): handle gemini not passing audio token usage data
* Chat Completions <-> Responses API Bridge Improvements (#11685)
* feat(anthropic/passthrough): pass dynamic api key/api base params to litellm.completion
allows calls to work with config.yaml
* fix(responses_api/transformation): fix passing dynamic params to responses api from .completion()
Allows responses api to work with config.yaml
* fix(langfuse.py): fix responses api usage logging to langfuse
* refactor(litellm_logging.py): add more generic solution for responses api usage logging
ensures it works across all logging integrations
* fix(litellm_logging.py): patch for anthropic messages not returning a pydantic object
it should ideally return a pydantic object, which would simplify checks and reduce errors
* fix(handler.py): correctly bubble up empty choices errors to litellm.completion
causes downstream errors as it is expected there is at least one choice set
* fix(response_metadata.py): allow model_info to be none
* fix(litellm_logging.py): copy object before mutating
* fix: fix lint check
* fix: fix linting error
* fix: fix linting error
---------
Co-authored-by: vanities <mischkeaa@gmail.com>
Co-authored-by: Cole McIntosh <82463175+colesmcintosh@users.noreply.github.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
* feat(anthropic/passthrough): pass dynamic api key/api base params to litellm.completion
allows calls to work with config.yaml
* fix(responses_api/transformation): fix passing dynamic params to responses api from .completion()
Allows responses api to work with config.yaml
* fix(langfuse.py): fix responses api usage logging to langfuse
* refactor(litellm_logging.py): add more generic solution for responses api usage logging
ensures it works across all logging integrations
* fix(litellm_logging.py): patch for anthropic messages not returning a pydantic object
it should ideally return a pydantic object, which would simplify checks and reduce errors
* fix(handler.py): correctly bubble up empty choices errors to litellm.completion
causes downstream errors as it is expected there is at least one choice set
* fix(response_metadata.py): allow model_info to be none
* fix(litellm_logging.py): copy object before mutating
* fix: fix lint check
- Updated the `_add_reasoning_system_prompt_if_needed` method to maintain the original format of list content when prepending the reasoning prompt.
- Adjusted tests to verify that both string and list content types are correctly handled, ensuring the reasoning prompt is added without altering the content structure.
- Updated the `_add_reasoning_system_prompt_if_needed` method to convert list content to strings before prepending the reasoning prompt.
- Adjusted tests to verify that system messages with list content are correctly transformed into strings, ensuring original content is preserved.
- Revised the reasoning support indicators in the Mistral model documentation for clarity.
- Improved the `_add_reasoning_system_prompt_if_needed` method to handle both string and list content types for system messages, ensuring the reasoning prompt is correctly prepended.
- Added a new test case to verify the functionality of adding the reasoning system prompt when the existing content is a list.
* fix(utils.py): convert stringified numbers to numbers
Closes https://github.com/BerriAI/litellm/issues/11266
* fix(convert_dict_to_model_response_object/): bubble up azure content_filter_results
* fix: fix linting error
* fix: fix linting errors
* fix(types/utils.py): ensure choices is correctly set
* fix: delete field if not set
* fix: expand scope of choicelogprobs value
* refactor(responses/): refactor to move responses_to_completion in separate folder
future work to support completion_to_responses bridge
allow calling codex mini via chat completions (and other endpoints)
* Revert "refactor(responses/): refactor to move responses_to_completion in separate folder"
This reverts commit ff87cb8958.
* feat: initial responses api bridge
write it like a custom llm - requires lesser 'new' components
* style: add __init__'s and bubble up the responses api bridge
* feat(responses/transformation): working sync completion -> responses and back bridge (non-streaming)
* feat(responses/): working async (non-streaming) completion <-> responses bridge
Allows calling codex mini via proxy
* feat(responses/): working sync + async streaming for base model response iterator
* fix: reduce function size
maintain <50 LOC
* fix(main.py): safely handle responses api model check
* fix: fix linting errors
* feat(parallel_request_limiter_v3.py): allows admin to enforce token rate limit based on just output tokens
Useful when trying to rate limit for primarily self hosted model use-cases
* test(test_parallel_request_limiter_v3.py): add unit test for token rate limit type
* feat(parallel_request_limiter_v3.py): return remaining token limits in header
* feat: return rate limit headers in response
* feat(parallel_request_limiter_v3.py): working rate limit response headers
* feat(parallel_request_limiter_v3.py): fix rate limit tracking for tpm when rpm also set
* feat(parallel_request_limiter_v3.py): show headers for key/user/team
* feat(parallel_request_limiter_v3.py): decrement max parallel request limiter on failure event
* feat(parallel_request_limiter_v3.py): add in-memory cache implementation of parallel request rate limiter
allows rate limiter to work even without redis cache setup
Work for GA of parallel request limiter v3
* refactor(proxy/hooks/__init__.py): replace with new parallel request handler
* test: update testing
* fix: fix ruff check
* fix: revert ga of multi instance rate limiting - needs more work to pass testing
* Added support for reasoning parameters in magistral models, including "reasoning_effort" and "thinking".
* Updated the MistralConfig class to handle reasoning system prompts.
* Implemented tests to verify reasoning functionality and ensure correct parameter mapping for magistral models.
* Enhanced the model prices JSON to reflect new reasoning capabilities.
* Checkpoint before follow-up message
* Add comprehensive tests for Deepgram transcription functionality
* clean up transform
* just use 1 test
* test cleanup
* test fix get_complete_url
* test rename file
* refactor deepgram URL construction
* add logging_obj.pre_call
* fix unused imports
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* fix(internal_user_endpoints.py): support user with `+` in email on user info
ensures user is correctly parsed from input
* fix(factory.py): support vertex function call args as None
handles empty string in args for vertex gemini calls
* docs(langfuse_integration.md): pin langfuse sdk version on docs
* fix(vertex_ai/): return empty dict, instead of none when empty string given
* refactor: reduce function size
* fix: fix linting errors
* fix: revert check
* fix(internal_user_endpoints.py): fix check
* test: update tests
* test: update tests
* docs(deploy.md): move docker recommendation to `main-stable`
* feat(enterprise/internal_user_endpoints.py): expose endpoint for checking available premium users
* feat(usage_indictor.tsx): add new element to help track remaining premium users
* feat(usage_indicator.tsx): show premium user remaining usage
allows users with user caps to know how much is left
* fix(vertex_and_google_ai_studio_gemini.py): bubble up stream is not finished, even if stop reason is given
prevents early completion of stream
Closes https://github.com/BerriAI/litellm/issues/11549
* fix(streaming_handler.py): respect is_finished = False in hidden params
internal logic for preventing ending stream early
* fix(litellm_license.py): add function to check if user is over limit
* fix(internal_user_endpoints.py): add function to check if user is over limit
* refactor: move test
* docs(customer_endpoints.py): document new param
* Feature/lasso guardrail (#9002)
* first version of lasso guardrail in litellm
* update to the new Lasso API
* change prod api_base and kill the request when lasso detect issue.
* change test for now api, local test pass
* add async tests
* all tests pass
* add docs for the new lasso guardrail
* Remove support for modes other than pre_call in Lasso guardrail
* code structure and naming
* only pre_call docs
* fix lint errors
* move test to the new location follows the same directory structure as litellm/.
* add lasso guard
* docs lasso docs
* add lasso guardrail
* fix lasso guardrail
---------
Co-authored-by: oroxenberg <oro@lasso.security>
* fix: fixes for transfer encoding error on aiohttp transport
* Update tests/test_litellm/llms/custom_httpx/test_aiohttp_transport.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>