Commit graph

23338 commits

Author SHA1 Message Date
Krrish Dholakia
8a1a90bc7a test: update test 2025-07-16 10:28:08 -07:00
Cole McIntosh
9c4b4303d3
fix: remove unused imports in completion_extras transformation (#12655)
- Remove unused GenericResponseOutputItem import
- Remove unused OutputFunctionToolCall import
- Fixes F401 linting errors
2025-07-16 10:23:13 -07:00
Krrish Dholakia
aa12a868c5 fix(migrate_keys.py): add script for migrating keys to new db 2025-07-16 10:18:36 -07:00
Cole McIntosh
d51aee1b84
Add GitHub Copilot LiteLLM tutorial (#12649)
* Add comprehensive GitHub Copilot + LiteLLM integration tutorial

- Complete setup guide from installation to production deployment
- Multiple configuration examples including authentication, load balancing, and cost tracking
- Docker and Kubernetes deployment configurations
- Troubleshooting section with common issues and solutions
- Best practices for security, monitoring, and reliability
- Usage examples for code completion, chat interface, and direct API integration

* Add concise GitHub Copilot + LiteLLM tutorial

- Create focused tutorial matching Gemini CLI style
- Step-by-step guide from installation to production deployment
- Multi-provider configuration examples (OpenAI, Anthropic, Bedrock)
- Load balancing and fallback configuration
- Docker deployment instructions
- Troubleshooting section with common issues
- Updated sidebar with clean title 'Use LiteLLM with GitHub Copilot'

* Refactor GitHub Copilot integration tutorial

- Removed outdated production deployment and direct API usage sections
- Streamlined troubleshooting steps for clarity
- Ensured documentation aligns with current best practices and configurations

* Add proper credit to Sergio Pino for GitHub Copilot tutorial

- Reference original DEV.to article in info box
- Add credits section acknowledging foundational work
- Maintain attribution to original author's guide
2025-07-16 09:40:27 -07:00
Krrish Dholakia
f4131b023e fix: don't fail request if unmapped item in responses list
not every responses item has a 1:1 mapping with chat completions
2025-07-16 09:25:26 -07:00
Ishaan Jaff
e8a748161f
[Bug Fix] grok-4 does not support the stop param (#12646)
* bug fix - using stop reason with grok 4

* fixes for XAI stop params

* test_xai_grok_4_stop_not_supported

* test_xai_grok_4_stop_not_supported
2025-07-16 09:19:25 -07:00
Krrish Dholakia
604075a36c test: update test 2025-07-16 09:15:05 -07:00
Krrish Dholakia
9cac629ca6 test: update test 2025-07-16 09:13:16 -07:00
Krrish Dholakia
446ed6039e docs(admin_ui_sso.md): document /fallback/login flow 2025-07-16 09:07:42 -07:00
Krrish Dholakia
5a8762b6a1 test: update tests 2025-07-16 08:56:41 -07:00
Cole McIntosh
c0935b9d58
Merge pull request #12648 from colesmcintosh/add-groq-moonshotai-kimi-k2-instruct
Add groq/moonshotai-kimi-k2-instruct model configuration
2025-07-16 09:50:52 -06:00
Cole McIntosh
ca40efab31 feat: add groq/moonshotai-kimi-k2-instruct model configuration
- Add model configuration for groq/moonshotai-kimi-k2-instruct
- Set max_tokens: 131072, max_input_tokens: 131072, max_output_tokens: 16384
- Configure pricing: 1e-06 input cost, 3e-06 output cost per token
- Enable function calling, response schema, reasoning, and tool choice support
2025-07-16 09:22:42 -06:00
Ishaan Jaff
0d1a1cfcfb
add together_ai/moonshotai/Kimi-K2-Instruct (#12645) 2025-07-16 07:53:27 -07:00
Michael Nguyen
f505977b62
Update bedrock nova micro and lite info (#12619) 2025-07-16 07:53:15 -07:00
Cole McIntosh
79e4d77bcf
fix: Handle circular references in spend tracking metadata JSON serialization (#12643)
* fix: Handle circular references in spend tracking metadata JSON serialization

- Fixes issue #12634 where circular references in metadata caused
  ValueError: Circular reference detected when logging spend data
- Adds _safe_json_dumps() function that detects and handles circular
  references by replacing them with placeholder strings
- Maintains full functionality for normal objects while preventing
  crashes from circular references
- Adds comprehensive tests for circular reference handling
- Critical fix for v1.74.3 stable release

* fix: Replace bare except clauses with specific Exception handling

- Fixes E722 linting errors in _safe_json_dumps function
- Maintains same error handling behavior while following best practices
- All tests continue to pass

* refactor: Use existing safe_dumps utility instead of custom implementation

- Replace custom _safe_json_dumps() with existing safe_dumps() from litellm_core_utils
- Remove duplicate code and leverage existing circular reference handling
- Update tests to use safe_dumps function
- Maintains same functionality while reducing code duplication
- All tests continue to pass
2025-07-16 07:24:41 -07:00
Krrish Dholakia
4e9440ee85 docs: cleanup docs 2025-07-16 07:18:13 -07:00
Krrish Dholakia
e22390a39a docs(openai.md): cleanup bridge doc 2025-07-15 22:59:10 -07:00
Krrish Dholakia
7064542504 docs(openai.md): document openai chat completions to responses api bridge 2025-07-15 22:50:23 -07:00
Krish Dholakia
1ce3558f96
fix(transformation.py): allows passing native responses api tools like web_search_preview, and mcp via .completion() (#12627)
Closes https://github.com/BerriAI/litellm/issues/12105
2025-07-15 22:45:42 -07:00
Krish Dholakia
955b504f7b
fix(proxy_server.py): fixes for handling team only models via `/v2/mo… (#12632)
* fix(proxy_server.py): fixes for handling team only models via `/v2/model/info`

ensures team only models show up on the correct team on `Models + Endpoints`

* test: update tests
2025-07-15 22:36:05 -07:00
Krish Dholakia
3ad4d9fc3e
fix(router.py): use more descriptive error message (#12629)
* fix(router.py): use more descriptive error message

* fix(proxy/_types.py): note `/team/member_update` is a self-managed route

route has it's own logic for rbac - enables team admins to update member permissions

Fixes issue where team admins on UI could not update member permissions

* fix(token_counter.py): move log line to being '.debug' instead of '.error'

Fixes https://github.com/BerriAI/litellm/issues/12269
2025-07-15 22:34:20 -07:00
Krrish Dholakia
85fe1d35e1 test: update test, remove old gemini models 2025-07-15 22:31:49 -07:00
Ishaan Jaff
84261f3ac8 test_create_delete_assistants 2025-07-15 21:35:25 -07:00
Ishaan Jaff
d8327b4740
[Bug Fix] [Bug]: Knowledge Base Call returning error (#12628)
* bug fix using vector stores as tools

* test_e2e_bedrock_knowledgebase_retrieval_with_llm_api_call_with_tools
2025-07-15 21:33:49 -07:00
yeahyung
782969bb05
(#11794) use upsert for managed object table rather than create to avoid UniqueViolationError (#11795)
* (#11794) use upsert for managed object table rather than create to avoid UniqueViolationError

* (#11794) use upsert for managed object table rather than create to avoid UniqueViolationError
2025-07-15 20:20:01 -07:00
Richard Tweed
197e7efa8f
fix: role chaining with webauthentication for aws bedrock (#12607)
* fix(bedrock): auto-generate session name when only aws_role_name is provided

Fixes #12583 - AWS role assumption not working correctly when aws_role_name
is provided without aws_session_name.

Previously, if only aws_role_name was provided in the config without
aws_session_name, the code would fall back to using environment credentials
instead of assuming the specified role. This was problematic in EKS/IRSA
environments where users want to assume a different role.

The fix:
- When aws_role_name is provided without aws_session_name, we now
  auto-generate a session name with format 'litellm-session-{timestamp}'
- This ensures role assumption happens as expected
- Added comprehensive test coverage for this scenario

* style: format test file with black

---------

Co-authored-by: openhands <openhands@all-hands.dev>
2025-07-15 20:17:09 -07:00
Daniele Scasciafratte
760ece5c8b
Update router.py (#12604) 2025-07-15 20:16:17 -07:00
Stefan Candra
685abf6871
Add token pricing for Together.ai Llama-4 and DeepSeek models (#12622)
* feat: add token pricing for Together.ai Llama-4 and DeepSeek models

* fix typo
2025-07-15 20:15:47 -07:00
Marc Abramowitz
40ccd2b70f
Add "keys import" command to CLI (#12620)
* Add "keys import" command to CLI

E.g.:

```
litellm-proxy keys import \
  --source-base-url=https://old-litellm.company.com \
  --source-api-key=$LITELLM_KEY \
  --dry-run
```

* Add --created-since option

* Add tests

* Fix lint errors

* Fix lint issues

* Fix lint errors

* Fix response.raise_for_status not being a thing

* Fix a mypy error
2025-07-15 20:14:43 -07:00
Stuart Geiger
e758b1b65f
rm claude instant 1 and 1.2 from model prices json (#12631) 2025-07-15 20:13:34 -07:00
Ishaan Jaff
2d6751a396
[Feat] MCP Gateway - allow using MCPs with all LLM APIs when using /responses with LiteLLM (#12546)
* add MCPResponsesAPIHelper

* rename LiteLLM_Proxy_MCP_Handler

* aresponses_api_with_mcp

* mock_responses_api_response

* test response with litellm proxy MCP

* add _should_use_litellm_mcp_gateway

* fix transform_mcp_tool_to_openai_responses_api_tool

* use correct _transform_mcp_tools_to_openai

* fix config.yaml

* fixes for native MCP handling

* docs MCP with litellm proxy

* aresponses_api_with_mcp

* fix linting

* fix mypy

* fix linting

* test_aresponses_api_with_mcp_mock_integration

* docs How it works when server_url="litellm_proxy"
2025-07-15 14:06:31 -07:00
Ishaan Jaff
d132f6e4e0
Troubleshooting (#12621) 2025-07-15 14:00:55 -07:00
Ishaan Jaff
46c950b17b
[Bug Fix] Add swagger docs for LiteLLM /chat/completions, /embeddings, /responses (#12618)
* update ProxyChatCompletionRequest

* use custom OpenAPI schema

* use customize_openapi_schema

* add_llm_api_request_schema_body

* add embeddings and response spec

* fix add_llm_api_request_schema_body

* TestCustomOpenAPISpec

* fixes linting

* fix linting

* bump timeout
2025-07-15 13:37:22 -07:00
Juan Carlos Moreno
a2f0bfc7e5
refactor(mcp): Make MCP_TOOL_PREFIX_SEPARATOR configurable from env (#12603)
The hardcoded hyphen (-) used as a tool name separator is now configurable via the MCP_TOOL_PREFIX_SEPARATOR environment variable.

- Replaced all hardcoded '-' instances with the MCP_TOOL_PREFIX_SEPARATOR constant.

- Updated error messages and docstrings to dynamically reference the new constant.

Tested and working fine.
Hope that helps!
2025-07-15 13:05:38 -07:00
Brian Caswell
45605f8362
add azure blob cache support (#12587)
* add support for Azure Blob caching

* add integration tests

* address feedback
2025-07-15 11:47:38 -07:00
tanjiro
720b94fd2b
Add Copy-on-Click for IDs (#12615)
* added the mcp_connect copy button to key_info_view

* user id copy button

* copy button on session view

* model id copy

* more copy buttons

* icon size
2025-07-15 11:25:57 -07:00
Ishaan Jaff
d2fe8894c8
[Bug Fix] Include /mcp in list of available routes on proxy (#12612)
* create GetRoutes helper class

* update get routes

* TestGetRoutes

* fix get_routes_for_mounted_app
2025-07-15 09:24:48 -07:00
Krrish Dholakia
e159ac932d docs: doc cleanup 2025-07-15 08:25:06 -07:00
Krrish Dholakia
0cbffe28a6 docs: doc cleanup 2025-07-15 08:23:42 -07:00
Krrish Dholakia
d850d1fa67 docs(controle_plane_and_data_plane.md): rename doc 2025-07-15 08:23:23 -07:00
Krrish Dholakia
45700b329b docs(bedrock.md): update doc
s
2025-07-15 07:28:00 -07:00
Dan McAulay
1b52815b70
fix(anthropic): fix streaming + response_format + tools bug (#12463)
* fix(anthropic): fix streaming + response_format + tools bug

- Fix _handle_json_mode_chunk to only convert response_format tools to content
- Regular user tools now remain as proper tool_calls in streaming mode
- Add comprehensive test for the fix
- Resolves issue where all tools were incorrectly converted to content chunks

Before: All tools converted to content with different indices
After: Only response_format tool converted, regular tools remain as tool_calls

* fix(anthropic): improve streaming + response_format + tools handling

* fix: lint error (too many statements)

* fix(anthropic): correct finish_reason for streaming response_format tools
2025-07-14 22:44:58 -07:00
Marcelo Díaz
094ce8f772
feat(gemini): Add custom TTL support for context caching (#9810) (#12541)
- Add ttl parameter to cache_control for Gemini models
- Support Google's TTL format (e.g., '3600s', '7200s')
- Implement robust TTL extraction and validation
- Extract TTL before system message transformation to handle all cases
- Add comprehensive test suite with 17 test cases in tests/test_litellm/
- Update documentation with TTL usage examples
- Maintain backward compatibility with existing cache_control usage

Fixes #9810
2025-07-14 22:30:54 -07:00
Anton
f05ec34e11
feat: Add envVars and extraEnvVars support to Helm migrations job (#12591)
- Add support for envVars (simple key-value pairs) in migrations job
- Add support for extraEnvVars (complex environment variable configurations)
- Include comprehensive test coverage for both envVars and extraEnvVars
- Ensure backward compatibility with existing configurations
- Tests verify proper rendering of environment variables in container spec
2025-07-14 22:24:13 -07:00
Krish Dholakia
58d9882102
refactor(prisma_migration.py): refactor to support use_prisma_migrate - for helm hook (#12600)
* refactor(prisma_migration.py): refactor to support use_prisma_migrate for helm hook

* fix(prisma_migration.py): don't use subprocess

* fix(prisma_migration.py): fix cli commands

* fix(prisma_migration.py): still run prisma generate
2025-07-14 22:11:46 -07:00
Krish Dholakia
49e9b73fcb
Claude 4 Bedrock /invoke route support + Bedrock application inference profile tool choice support (#12599)
* docs(config_settings.md): document enable_json_schema_validation

Closes https://github.com/BerriAI/litellm/issues/12518

* fix(utils.py): add claude-sonnet-4 on bedrock support

Fixes https://github.com/BerriAI/litellm/issues/12366

* refactor(utils.py): move list to getter in function

more maintainable

* fix(utils.py): handle bedrock_converse in provider check

Fixes https://github.com/BerriAI/litellm/issues/11751
2025-07-14 21:42:25 -07:00
Krish Dholakia
7c392475e6
Control Plane + Data Plane support (#12601)
* feat(route_checks.py): allow admin to disable proxy management endpoints on instance

useful for preventing multiple instances from doing admin actions

* docs(scaling_multiple_instances.md): add architecture doc on scaling multiple litellm instances

provide guidance on scaling proxy

* docs(scaling_multiple_instances.md): add doc on scaling across multiple regions for litellm

* fix(route_checks.py): allow disabling llm api endpoints on an instance

allows pure admin instance to exist

* refactor(enterprise/route_checks.py): refactor env var checks

* refactor: finish refactoring

* docs(control_plane_and_data_plane.md): refactor docs

* test: update tests
2025-07-14 21:31:56 -07:00
Ishaan Jaff
a91e115f9a bump: version 1.74.3 → 1.74.4 2025-07-14 20:06:50 -07:00
Ishaan Jaff
57b0b4edf3
[Bug fix] [Bug]: Verbose log is enabled by default (#12596)
* test find_set_verbose_assignments

* fix set verbose

* test set verbose

* fix unused import
2025-07-14 20:06:16 -07:00
tanjiro
c0f1b8119e
Wildcard model filter (#12597)
* filter by wildcard models

* merge conflict
2025-07-14 20:01:10 -07:00