Commit graph

29901 commits

Author SHA1 Message Date
Yuneng Jiang
0c560f70b7
chore: fixes 2026-04-04 23:13:38 -07:00
Cursor Agent
e317d20573 feat: Redact messages in spend logs' proxy_server_request
This change ensures that sensitive message content within the `proxy_server_request` field of spend logs is redacted when message logging is turned off or when the redaction header is used. This enhances privacy and security by preventing the logging of potentially sensitive user inputs in the request details.

Co-authored-by: ishaan <ishaan@berri.ai>
2026-01-09 22:06:34 +00:00
Sameer Kankute
0c7db97ad5
Merge pull request #18871 from BerriAI/litellm_fix_test_count_tokens_caching
Fix :test_count_tokens_caching
2026-01-10 01:20:55 +05:30
Harshit Jain
8a683d9a6a
Add fix for bedrock_cache, metadata and max_model_budget (#18872) 2026-01-10 01:09:00 +05:30
Sameer Kankute
aba7dcea9c Fix : litellm import error 2026-01-10 01:03:50 +05:30
Sameer Kankute
777ae4f530 Fix :test_count_tokens_caching 2026-01-10 00:57:11 +05:30
Robin
0575bd2d1c
feat: update prices json for novita provider (#18540)
* feat: add novita models

* feat: ci

* feat: add novita support josn
2026-01-10 00:48:06 +05:30
Shivam Rawat
691505cdae
added fix for org level budget enforcement (#18813) 2026-01-10 00:44:40 +05:30
Shivam Rawat
43dd0e6ef5
remove model before casting it in the transformation (#18810) 2026-01-10 00:43:38 +05:30
Cesar Garcia
c19c97591e
fix: align max_tokens with max_output_tokens for consistency (#18820)
* fix: align max_tokens with max_output_tokens for consistency

Fixed inconsistent max_tokens definitions in model_prices_and_context_window.json.
According to LiteLLM convention, max_tokens should equal max_output_tokens when available.

Models fixed:
- deepseek-chat: 131072 → 8192 (now equals max_output_tokens)
- dashscope/qwen-flash: 1000000 → 32768 (now equals max_output_tokens)
- databricks/databricks-gemma-3-12b: 128000 → 32000 (now equals max_output_tokens)

This ensures consistency across all providers where max_tokens represents
the maximum number of tokens that can be generated in the output.

* fix: align max_tokens with max_output_tokens for 244 models

- Fix 244 models where max_tokens != max_output_tokens
- Add test to validate max_tokens consistency and prevent regressions

According to model_prices_and_context_window.json spec:
- max_tokens is a LEGACY parameter
- Should always equal max_output_tokens when both are present

This ensures consistency across all model definitions.
2026-01-10 00:37:45 +05:30
DominikHallab
fa2b0fb533
docs: Update header to be markdown bold by removing space (#18846) 2026-01-10 00:32:21 +05:30
mel2oo
7161f41746
Fix: google_genai streaming adapter provider handling (#18845) 2026-01-10 00:31:52 +05:30
Alexsander Hamir
3a22fa89c4
Refactor ProviderConfigManager.get_provider_chat_config for O(1) performance (#18867) 2026-01-09 10:40:06 -08:00
Martin Gauthier
dd087c8cca
🐛 fix: propagate headers in router embedding calls (#18844)
Fix router embedding methods to properly propagate proxy model
configuration headers to LLM API calls by calling
_update_kwargs_before_fallbacks() just like completion() does.

Previously, router.embedding() and router.aembedding() manually
set num_retries and metadata but didn't call
_update_kwargs_before_fallbacks(), which meant default_litellm_params
(including headers) were not propagated correctly.

Changes:
- Replace manual kwargs setup with _update_kwargs_before_fallbacks()
  in _embedding method (litellm/router.py:3318)
- Apply Black formatting to router.py for consistency
- Add comprehensive unit tests for header propagation
- Add integration tests for various router configurations

Tests verify:
- Headers from default_litellm_params are included in embedding calls
- Metadata (model_group) is properly set
- Consistency between completion() and embedding() behavior
- Support for deployment-specific headers, fallbacks, and retries
2026-01-09 23:57:18 +05:30
Sameer Kankute
8dac83e093
Merge pull request #18859 from BerriAI/litellm_azure_image_gen_fix
Fix: response_format leaking into extra_body
2026-01-09 23:09:27 +05:30
Andrés
9768eca33e
fix(azure): add logprobs support for Azure OpenAI GPT-5.2 model (#18856)
* fix(azure): add logprobs support for Azure OpenAI GPT-5 models

Azure OpenAI GPT-5 models (including gpt-5.2) support logprobs
parameters, unlike OpenAI's GPT-5 reasoning models. This fix
overrides the parent class restriction to enable logprobs for Azure.

Changes:
- Override get_supported_openai_params() in AzureOpenAIGPT5Config
- Add "logprobs" and "top_logprobs" to supported params
- Add comprehensive tests for logprobs functionality

Testing:
- Verified with direct Azure API calls to gpt-5.2
- API version: 2025-01-01-preview
- Successfully returns logprobs data

Related: #7974, #4022

* refactor: restrict logprobs support to gpt-5.2 only

Only gpt-5.2 has been verified to support logprobs on Azure.
Other gpt-5 variants (gpt-5, gpt-5.1) have not been tested.

Changes:
- Add conditional check for is_model_gpt_5_2_model()
- Update tests to be specific to gpt-5.2
- Add negative tests for gpt-5 and gpt-5.1
- Update documentation to reflect gpt-5.2 specificity
2026-01-09 22:57:50 +05:30
Cesar Garcia
153fd7ad0a
fix: prevent Prisma migration workflow from running in forks (#18863)
- Add repository check to only run workflow in BerriAI/litellm
- Prevents workflow failures and wasted resources in forked repositories
- Avoids confusion for external contributors
2026-01-09 22:56:32 +05:30
yuneng-jiang
c3c67d445b
Merge pull request #18849 from BerriAI/litellm_e2e_page_routing
[Infra] UI - E2E Test: Refactor Page Settings + Test for Page Navigation
2026-01-09 09:07:25 -08:00
yuneng-jiang
68afc1276e
Merge pull request #18848 from BerriAI/litellm_ui_unit_test_coverage_006
[Infra] UI - Unit Test: Adding Tests to Expand Coverage
2026-01-09 09:07:16 -08:00
Harshit Jain
819468554f
fix(security): prevent expired key plaintext leak in error response (#18860) 2026-01-09 22:27:39 +05:30
Sameer Kankute
c0b05fc47a
Merge pull request #18250 from sjmatta/claude/fix-issue-17910-QBgDq
[Fix] Nova model detection for Bedrock provider (#17910)
2026-01-09 17:30:51 +05:30
Sameer Kankute
be28fcd463
Merge pull request #18858 from raghav-stripe/raghav-add-bedrock-tokencounter
feat: add Bedrock as a backend API for token counting
2026-01-09 17:10:52 +05:30
Sameer Kankute
072e3dcd0c
Merge pull request #18715 from BerriAI/litellm_staging_01_06_2026
Staging 01/06/2026
2026-01-09 17:10:04 +05:30
Sameer Kankute
e7efd51bd7
Merge branch 'main' into litellm_staging_01_06_2026 2026-01-09 17:06:17 +05:30
Sameer Kankute
bb9347207b
Merge pull request #18833 from BerriAI/litellm_staging_01_08_2026
Litellm staging 01 08 2026
2026-01-09 17:04:29 +05:30
Sameer Kankute
844c766c65
Merge pull request #18763 from BerriAI/litellm_staging_01_07_2026
Staging - 01/07/2026
2026-01-09 17:01:58 +05:30
Sameer Kankute
ffa0d6706c Fix: response_format leaking into extra_body 2026-01-09 16:53:35 +05:30
Sameer Kankute
d67b475ba4
Merge pull request #18853 from jutaz/bugfix/unbound-local-error
Fix: Add thought_signatures to VertexGeminiConfig and test
2026-01-09 16:27:12 +05:30
Raghav Jhavar
f5ee89dcf6 fix broken imports 2026-01-09 17:43:23 +07:00
Raghav Jhavar
0bdc411fb7 remove comment 2026-01-09 17:24:33 +07:00
Raghav Jhavar
ba78194ff1 add support for bedrock in token counting api 2026-01-09 17:08:07 +07:00
Sameer Kankute
940138f585
Merge pull request #18629 from akraines/fix/wildcard-routing-errors
fix: Improve error messages and validation for wildcard routing with multiple credentials
2026-01-09 14:51:13 +05:30
Justas Brazauskas
c0ee5da444
Fix: Add thought_signatures to VertexGeminiConfig and test 2026-01-09 10:03:45 +02:00
YutaSaito
0e67fdaef0
Merge pull request #18850 from BerriAI/litellm_feat_mcp-registry
[feat] add mcp registry
2026-01-09 16:16:46 +09:00
Yuta Saito
022db6c9ed feat: add mcp registry 2026-01-09 15:07:39 +09:00
yuneng-jiang
4d28359608 e2e refactoring 2026-01-08 21:50:19 -08:00
yuneng-jiang
bcbf8d1de3 Adding unit test to expand unit testing coverage 2026-01-08 20:23:54 -08:00
Sameer Kankute
d9b275e62a
Merge pull request #18806 from BerriAI/litellm_vertex_ai_api_key_support
[FEAT]: Add support for Vertex AI API keys
2026-01-09 09:44:36 +05:30
YutaSaito
c81ac6cf46
Merge pull request #18841 from BerriAI/litellm_fix_cloudzero_db
[fix] how to execute cloudzero sql
2026-01-09 12:37:07 +09:00
Yuta Saito
48bc5ccb4f fix: test 2026-01-09 11:23:47 +09:00
YutaSaito
e678540843
Merge pull request #18836 from drorIvry/main
feat: added qualifire eval webhook
2026-01-09 11:06:45 +09:00
Yuta Saito
f129f598a0 fix: how to execute cloudzero sql 2026-01-09 10:52:48 +09:00
YutaSaito
c27bfddca0
Merge pull request #18837 from BerriAI/litellm_docs_focus
[docs] add focus
2026-01-09 08:00:20 +09:00
Yuta Saito
c38294dc16 docs: add focus 2026-01-09 07:55:55 +09:00
YutaSaito
661f03058c
Merge pull request #18802 from BerriAI/litellm_feat_focus_backend
[feat] Focus export support
2026-01-09 07:21:04 +09:00
Dror Ivry
e6c41c8f47
docs 2026-01-09 00:08:47 +02:00
Dror Ivry
5793f8b866
docs 2026-01-09 00:08:03 +02:00
Dror Ivry
8ff4cb89e1
feat: added qualifire eval webhook 2026-01-09 00:03:26 +02:00
Yuta Saito
844577f2d4 refactor: lazy-load focus export engine and isolate polars deps 2026-01-09 06:39:24 +09:00
yuneng-jiang
1d19974737
Merge pull request #18835 from BerriAI/litellm_ui_build_fix_00
[Infra] Fixing UI Build
2026-01-08 13:12:56 -08:00