Commit graph

1447 commits

Author SHA1 Message Date
katsuhiro muto
d5e686b3e8
[Fix] Support service_tier in chat completion (#15693)
* Support service_tier

* fix test
2025-10-18 13:55:54 -07:00
Sameer Kankute
282eedb72b Add responses mode to health check 2025-10-17 22:06:29 +05:30
Krish Dholakia
3ff073c811
UI - add arize on ui, LLMs - clarifai refactor to openai compatible route, added azure ai/grok-4 model family
* added oauth mcp to docs

* added azure ai/grok-4 model family

* Revert "added oauth mcp to docs"

This reverts commit 950b7cef44.

* fix: arize ui integration

* need to remove a file

This reverts commit d6c877b73a.

* fix: add arize from ui

* updated clarifai functions to openai compatible (#15615)

* fix: npm build errors

* Snowflake provider support: added embeddings, PAT, account_id (#15372)

* snowflake support PAT, account_id and embeddings

* format

* test embeddings

* format

* complete test

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>

* Revert "Snowflake provider support: added embeddings, PAT, account_id (#15372)" (#15632)

This reverts commit c6d58e5b4a.

---------

Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Mubashir Osmani <ilikewafflesomcuh@gmail.com>
Co-authored-by: mogith-pn <143642606+mogith-pn@users.noreply.github.com>
Co-authored-by: Andrey <elkin.andr@gmail.com>
2025-10-16 20:39:15 -07:00
TensorNull
55a6dd3a8b feat(cometapi): Add CometAPI provider support (embeddings, image generation, docs)
- Add CometAPI embedding and image generation transformations and configs
- Add image cost calculator and export/init files
- Register provider in constants, utils, main (embedding path) and sidebars
- Add CometAPI docs page and cookbook notebook (Colab) for usage examples
2025-10-16 13:08:14 +08:00
Alexsander Hamir
2c9356c437
[Fix] - shared session parsing and usage issue (#15440)
* Fix: Add shared_session to all_litellm_params to prevent JSON serialization error

The shared_session parameter (aiohttp.ClientSession) was being passed through
to provider API calls, causing "Object of type ClientSession is not JSON
serializable" errors during embedding requests.

Added shared_session to the all_litellm_params list so it's properly filtered
out as a LiteLLM-internal parameter and not passed to the provider's API.

* Fix: Add shared_session support for embedding calls with connection pooling

The shared_session parameter was not being properly handled in embedding calls,
causing it to be passed through to provider API requests where it's not needed.

Changes:
- Added shared_session to all_litellm_params to filter it from provider API request body
- Extract shared_session in main embedding() function and pass it explicitly
- Updated OpenAI embedding handlers (embedding() and aembedding()) to accept shared_session
- Pass shared_session to _get_openai_client for HTTP client creation

This enables proper connection pooling for embedding requests when shared_session
is provided, improving performance for high-throughput scenarios.

* test: add regression test for shared_session in embedding calls

Add comprehensive test to prevent JSON serialization errors when using
shared_session.

The test verifies two critical aspects:
1. shared_session is in all_litellm_params to prevent "Object of type
   ClientSession is not JSON serializable" errors
2. shared_session flows through the complete call chain across 6 layers:
   - litellm.embedding()
   - OpenAI.embedding/aembedding()
   - _get_openai_client()
   - AsyncHTTPHandler.create_client()
   - _create_async_transport()
   - _create_aiohttp_transport()

Similar to test_acompletion_session_reuse_e2e.py but focused on
embedding endpoints. Uses inspect.getsource() to verify the parameter
is not only accepted but actually passed through each layer.
2025-10-10 19:26:53 -07:00
Ishaan Jaff
52bbabd788
[Feat] Support for Vertex AI Gemma Models on Custom Endpoints (#15397)
* TestVertexGemmaiCompletion

* test vertex Gemma

* fix file name

* fix file naming

* add VertexAIGemmaModels

* add cost_router for vertexai

* fix main.py

* fix VertexGemmaConfig

* fix Vertex AI Gemma-AI Models Handler

* docs gemma

* fix ids

* test fix

* ruff check fixes

* docs fix

* docs fix
2025-10-09 19:20:02 -07:00
Sameer Kankute
cd25782359
fix passing headers for gemini (#15231) 2025-10-06 12:59:19 -07:00
Henry H Wang
7c4439ba0a
Merge branch 'BerriAI:main' into gemini-adapter-fixes 2025-10-01 09:53:51 +08:00
Eddie Richter
1e1e4c36ac Fixing key name 2025-09-30 12:12:24 -06:00
Eddie Richter
ae92404d05 Initial addition of Lemonade provider. 2025-09-30 12:12:24 -06:00
Henry Wang
33218606b8 fix mypy check issues 2025-09-30 10:37:21 +08:00
Henry Wang
0c1104fbe0 fix(lint): Resolve F821 Undefined name errors in litellm/main.py 2025-09-29 22:42:09 +08:00
jacobi petrucciani
6249f4efa5 fix passthrough of atranscription into kwargs going to upstream provider 2025-09-28 12:46:43 -04:00
Henry Wang
b46407fa76 feat(gemini): Add full support for native Gemini API translation
This commit implements a complete, end-to-end fix for the native Gemini API translation feature, allowing requests to be correctly routed to other model providers via `model_group_alias`.

The original implementation was broken, causing `systemInstruction` and `tools` to be dropped from requests. This was resolved by refactoring the Gemini endpoint to use a dedicated translation path, similar to the Anthropic adapter.

Additionally, this commit hardens the streaming response adapter to correctly handle tool calls generated by the newly-fixed request path. Key improvements to the response handling include:

- Replaced the fragile `id`-based tool call tracking with a robust `index`-based accumulation logic.
- Fixed a memory leak and improved logging in the stream finalization process.
- Prevented empty, non-compliant chunks from being sent to the client during tool call streaming.
- Optimized the accumulator to skip and log superfluous empty chunks sent by some models.
2025-09-28 21:01:31 +08:00
Krrish Dholakia
575e360a60 fix: fix error on main 2025-09-27 13:31:11 -07:00
Krrish Dholakia
e02e58971f fix: fix linting errors on main 2025-09-27 13:06:05 -07:00
Krrish Dholakia
5cbfa42bb4 fix: fix linting error 2025-09-27 13:02:32 -07:00
Krrish Dholakia
3bf3a7fa83 fix: fix linting errors on main.py 2025-09-27 12:53:42 -07:00
Krrish Dholakia
f00a32d04c fix: fix linting errors 2025-09-27 12:47:13 -07:00
Alexsander Hamir
eaa04cd8ce
fix: use fastuuid helper (#14903)
* fix: use fastuuid helper across the codebase

First batch of changes, simple drop in replacement.

* second batch of changes

* fixed: script mistake on helper file
2025-09-25 15:47:01 -07:00
Krish Dholakia
64a305f9dd
Merge pull request #14721 from dharamendrak/reuse-aiohttp-session-http-handler
feat: Add shared_session parameter for aiohttp ClientSession reuse
2025-09-23 18:07:32 -07:00
Ishaan Jaffer
724d8b5afa fix linting 2025-09-23 14:01:48 -07:00
Ishaan Jaffer
426c6fea9f Revert "Merge pull request #14761 from uzaxirr/feat/sdk-additional-headers"
This reverts commit 8628c265b9, reversing
changes made to be193fbffd.
2025-09-23 13:59:54 -07:00
Dharamendra Kumar
34f51a20eb Merge branch 'main' into reuse-aiohttp-session-http-handler 2025-09-23 10:36:35 -07:00
Anubhav Singh
fde4dcb0b8
Merge branch 'main' into wandb-inference 2025-09-23 00:14:11 +05:30
uzaxirr
d8d5581f29 Resolve merge conflicts with main branch
- Accept main branch's consolidation of Cohere providers
- Preserve header implementation with proper three-tier merging
- Replace outdated header handling in Cohere sections with consolidated approach
- Maintain backward compatibility and functionality
2025-09-22 05:59:28 +05:30
uzaxirr
50d717cde0 Apply Black formatting and fix Ruff issues
- Format code with Black to meet style requirements
- Fix auto-fixable Ruff linting issues
- Maintain header implementation functionality
2025-09-22 04:46:08 +05:30
uzaxirr
3986b073c9 feat: Add SDK support for additional headers 2025-09-21 14:58:14 +05:30
Dharamendra Kumar
77a39e7ca9 feat: Add shared_session parameter for aiohttp ClientSession reuse
Allow passing aiohttp.ClientSession to acompletion() calls for better
performance and resource management. Includes debug logging, tests,
and documentation. Backward compatible.
2025-09-19 01:46:53 -07:00
Ishaan Jaff
759bb6cff5
fix: cohere generate api is deprecated (#14676) 2025-09-18 07:27:51 -07:00
Krish Dholakia
bf0dd4a284
Merge pull request #14418 from iabhi4/deep-copy-issue
fix: avoid deepcopy crash with non-pickleables in Gemini/Vertex
2025-09-16 22:55:31 -07:00
Anubhav Singh
a9667e5930
Merge branch 'BerriAI:main' into wandb-inference 2025-09-16 16:46:58 +05:30
Tim Elfrink
afd720a62f Fix CompactifAI provider tests and implementation
- Add missing provider_config parameter in main.py for proper HTTP handler integration
- Update tests to use correct respx mocking pattern with litellm.disable_aiohttp_transport
- Add get_error_class method to CompactifAI transformation for proper error handling
- Fix authentication error test to expect APIConnectionError instead of AuthenticationError
- All 8 CompactifAI tests now pass successfully
2025-09-15 22:03:42 +02:00
Anubhav Singh
67276a8151
Merge branch 'main' into wandb-inference 2025-09-15 18:34:16 +05:30
Tim Elfrink
9521414efa Resolve merge conflict by including both CompactifAI and OVHCloud providers
- Keep CompactifAI provider detection logic
- Include new OVHCloud provider from main branch
- Both providers now work correctly with model prefix detection
2025-09-14 23:03:18 +02:00
Krish Dholakia
56fd60b140
Merge pull request #14494 from eliasto/feat/ovhcloud-ai-edpoints-provider
feat: Add OVHCloud AI Endpoints as a provider
2025-09-14 00:45:08 -07:00
Krrish Dholakia
8443000ca4 fix(main.py): route vllm calls via the openai sdk route
consistent with other openai-like implementations
2025-09-13 11:49:05 -07:00
Krish Dholakia
269515e525
Merge branch 'main' into litellm_dev_09_12_2025_p1 2025-09-13 10:10:30 -07:00
Tim Elfrink
9402dc35aa Integrate CompactifAI provider into LiteLLM core
- Add COMPACTIFAI to LlmProviders enum for type safety
- Register CompactifAIChatConfig in ProviderConfigManager
- Import CompactifAIChatConfig in main __init__.py
- Add 'compactifai/' model prefix detection in get_llm_provider()
- Wire CompactifAI completion handler in main.py routing logic
- Support COMPACTIFAI_API_KEY environment variable
- Enable base_llm_http_handler for OpenAI-compatible requests
- Maintain consistency with existing provider integration patterns
2025-09-13 08:42:04 +02:00
Krrish Dholakia
82091de393 feat(hosted_vllm/): transcription endpoint support
Closes https://github.com/BerriAI/litellm/issues/361#issuecomment-3244548055
2025-09-12 17:15:14 -07:00
Elias TOURNEUX
ef9d1ddc40
feat: Add OVHCloud AI Endpoints as a provider 2025-09-12 13:37:03 +02:00
Krrish Dholakia
d89152bb2c feat(litellm_logging.py): support new litellm debug parameter - litellm_request_debug on requests
enables printing raw request when flag is set to true on requests
2025-09-11 20:04:24 -07:00
iabhi4
384ad7e99c fix: avoid deepcopy crash with non-pickleables in Gemini/Vertex 2025-09-10 12:59:03 -07:00
xprilion
667481e75b (feat): Add W&B Inference to LiteLLM 2025-09-11 00:07:30 +05:30
Derek Worthen
59360e64f9 Fix embeddings using azure_ad_token_provider. 2025-09-09 07:01:05 -07:00
Krish Dholakia
b9ce3a1587
Merge pull request #12416 from dotmobo/feature/fix-alloy
feat: add a health_check_voice parameter in model_info
2025-09-08 23:12:48 -07:00
Krish Dholakia
ba10173ec7
Merge branch 'main' into heroku-llms 2025-09-06 22:10:20 -07:00
Krish Dholakia
f67339a86c
Merge pull request #14028 from onlylhf/volcengine-embedding-support
Add Volcengine embedding module with handler and transformation logic
2025-09-04 21:01:25 -07:00
katsuhiro muto
ca43514db4
[Feat] Support reasoning_effort in Groq (#14207)
* Support reasoning_effort in groq

* add test
2025-09-03 10:43:47 -07:00
Sameer Kankute
4adfd18bc6
[Feat]Add support for safety_identifier parameter in chat.completions.create (#14174)
* Add support for safety_identifier parameter in chat.completions.create

* make sure param is getting actually passed to the raw api
2025-09-02 09:37:08 -07:00