Commit graph

30138 commits

Author SHA1 Message Date
Ishaan Jaffer
190b932ded docs fix 2026-01-13 19:02:19 -08:00
Ishaan Jaffer
8f28a60f98 add SDK level fixes 2026-01-13 19:00:09 -08:00
Ishaan Jaffer
b6dcaf0cf6 docs fix 2026-01-13 18:57:16 -08:00
Ishaan Jaffer
d105580fba doc fix 2026-01-13 18:54:39 -08:00
Ishaan Jaffer
c39492fa58 fix 2026-01-13 18:52:50 -08:00
Ishaan Jaffer
3322f83bdf docs fix 2026-01-13 18:50:27 -08:00
Ishaan Jaffer
d3135cc5ce fix rate limiting 2026-01-13 18:49:05 -08:00
Ishaan Jaffer
d6ce6377bd docs fix 2026-01-13 18:43:41 -08:00
Ishaan Jaffer
aedf648bc2 fix 2026-01-13 18:42:09 -08:00
Ishaan Jaffer
ad040d4981 fix 2026-01-13 18:38:14 -08:00
Ishaan Jaffer
22eecc0678 bring back older diagram 2026-01-13 18:36:36 -08:00
Ishaan Jaffer
73b4107b9c remove bloat 2026-01-13 18:33:41 -08:00
Ishaan Jaffer
7ec0c096d5 v0 of this 2026-01-13 18:31:19 -08:00
YutaSaito
a66e007574
Merge pull request #19051 from BerriAI/litellm_fix_mcp-rest-auth-checks
[fix] mcp rest auth checks
2026-01-14 10:39:28 +09:00
Alexsander Hamir
15c3bc219b
[Refactor] Add CI enforcement for O(1) operations in _get_model_cost_key to prevent performance regressions (#19052)
* Optimize _get_model_cost_key to avoid expensive scans

- Remove expensive O(n) scan fallback that was causing 42.87% CPU overhead
- Only scan when size mismatch detected (O(1) check)
- Add warning in docstring: Only O(1) lookup operations are acceptable
- Clean up comments to be more concise
- Keep stale entry rebuild for pop() case (only triggers when stale entry found)

This fixes the performance issue where the scan was being triggered on every
failed lookup, causing severe CPU overhead during router operations.

* Add code quality check to enforce O(1) operations in _get_model_cost_key

- Add check_get_model_cost_key_performance.py to statically analyze _get_model_cost_key
- Detects O(n) operations (loops, comprehensions, problematic function calls)
- Recursively checks called functions to find nested O(n) operations
- Allows conditional O(n) rebuilds in helper functions (_rebuild_model_cost_lowercase_map, _handle_stale_map_entry_rebuild, _handle_new_key_with_scan)

* Integrate _get_model_cost_key performance check into CI pipeline

- Add check_get_model_cost_key_performance.py to check_code_and_doc_quality job
- Ensures O(1) requirement is enforced in CI to prevent performance regressions

* Remove unused performance test and clean up utils.py

- Remove test_get_model_info_performance.py (no longer needed)
- Remove extra blank line in utils.py

* Document allowed helper functions and exception process in _get_model_cost_key

- Add documentation listing allowed helper functions with O(n) operations
- Explain why these are acceptable (conditionally called)
- Add instructions for adding new exceptions to check_get_model_cost_key_performance.py

* Fix docstring detection and type checker error in performance check

- Add proper docstring tracking to skip docstring content (fixes false positive for 'map' in docstring)
- Add None check for docstring_quote to fix type checker error
- Restore _handle_new_key_with_scan to allowed_helpers list

* Remove check_get_model_cost_key_performance from CI pipeline

- Temporarily remove the performance check from CI to avoid blocking builds

* Restore performance check and remove memory leak tests from CI

- Add back check_get_model_cost_key_performance.py to CI pipeline
- Remove memory_leak_tests job that was causing port conflicts

* Remove extra blank line in CI config
2026-01-13 17:08:03 -08:00
yuneng-jiang
2ea6fcb584
Merge pull request #19050 from BerriAI/litellm_ui_usage_top_x
[Feature] UI - Usage: Allow Top Virtual Keys and Models to Show More Entries
2026-01-13 17:06:07 -08:00
Raghav Jhavar
272a48d880
[bug fix] do not fallback to token counter if disable_token_counter is enabled (#19041)
* do not fallback to token counter if disable_token_counter is enabled, and return errors instead

* add exceptions and exception utils to map the same as /v1/chat/completions

* use safe_json_loads
2026-01-13 16:53:38 -08:00
Yuta Saito
b54b12d0cc fix: lint 2026-01-14 09:53:25 +09:00
yuneng-jiang
70987ffdc8
Merge pull request #19047 from BerriAI/litellm_usage_column_names
[Fix] UI - Usage: Team ID and Team Name in Export Report
2026-01-13 16:26:26 -08:00
yuneng-jiang
212eec5a6c
Merge pull request #18997 from BerriAI/litellm_ui_key_generate_readable
[Feature] UI - Simplify Key Generate Permission Error
2026-01-13 16:26:15 -08:00
yuneng-jiang
52d5eccf3d
Merge pull request #19010 from BerriAI/litellm_ui_filter_components
[Refactor] UI - User and Team Table Filters to Reusable Component
2026-01-13 16:26:07 -08:00
yuneng-jiang
e6ca97a9a5 top model show N 2026-01-13 16:21:14 -08:00
yuneng-jiang
ddf2b24901 Top virtual keys show N select 2026-01-13 15:50:34 -08:00
Alexsander Hamir
a1dd3ead4d
[Perf] Remove bottleneck causing high CPU usage & overhead under heavy load (#19049) 2026-01-13 15:22:09 -08:00
Yuta Saito
df37770a70 test: add permission test 2026-01-14 07:58:21 +09:00
yuneng-jiang
308d24d067 Resolve Team ID and Team name in export 2026-01-13 14:42:28 -08:00
Cesar Garcia
a445d8a4b4
fix(pricing): correct cache_read pricing for gemini-2.5-pro models (#18157)
- Fix cache_read_input_token_cost: 3.125e-07 → 1.25e-07 ($0.125/1M)
- Add cache_read_input_token_cost_above_200k_tokens: 2.5e-07 ($0.25/1M)

Models updated:
- gemini-2.5-pro
- gemini-2.5-pro-exp-03-25
- gemini-2.5-pro-preview-03-25
- gemini-2.5-pro-preview-05-06
- gemini-2.5-pro-preview-06-05
- gemini-2.5-pro-preview-tts
- gemini/gemini-2.5-pro
- gemini/gemini-2.5-pro-preview-03-25
- gemini/gemini-2.5-pro-preview-05-06
- gemini/gemini-2.5-pro-preview-06-05
- gemini/gemini-2.5-pro-preview-tts

Pricing source: https://ai.google.dev/gemini-api/docs/pricing
2026-01-14 04:09:24 +05:30
nulone
478bdcb60b
fix(model_prices): sync DeepSeek chat/reasoner to V3.2 pricing (#18884) 2026-01-14 03:52:50 +05:30
Cesar Garcia
d03c5017ff
fix: correct context window sizes for GPT-5 model variants (#18928)
* fix: correct context window sizes for GPT-5 model variants

Updates max_input_tokens for GPT-5, GPT-5.1, and GPT-5.2 model variants
to match OpenAI's official specifications, resolving issue #18927.

Changes:
- GPT-5.1 (base, codex variants): 272k → 400k tokens
- GPT-5.1-chat variants: 272k → 128k tokens (with max_output 16,384)
- GPT-5 (base): 272k → 400k tokens
- GPT-5-chat: 272k → 128k tokens (with max_output 16,384)
- GPT-5-codex: 272k → 400k tokens
- GPT-5-mini: 272k → 400k tokens
- GPT-5-nano: 272k → 400k tokens
- GPT-5-pro: 272k → 400k tokens (max_output 272k)
- GPT-5 dated versions (2025-08-07): 272k → 400k tokens

Affected providers: OpenAI, Azure (all regions), OpenRouter

Fixes #18927

* fix: correct Azure GPT-5 context window limits to match Azure docs

Azure OpenAI has different limits than OpenAI for GPT-5 models.

Changes:
- Azure GPT-5 models: max_input_tokens 400k → 272k (Azure limit)
- Azure GPT-5 Pro: max_output_tokens 272k → 128k (Azure limit)
- OpenAI GPT-5 models: remain at 400k (correct)
- OpenRouter models: remain at 400k (routes to OpenAI)

Azure docs specify 272k input + 128k output = 400k total context.
OpenAI allows full 400k input + 128k output.

* fix: correct Azure GPT-5 max_input_tokens to 272k

Azure has explicit input limit of 272k tokens (not 400k like OpenAI).
Context window 400k = 272k input + 128k output for Azure.
OpenAI allows flexible input up to 400k (context - output).
2026-01-14 03:49:47 +05:30
Robin
b7c5662273
Fix: update novita models prices (#19005)
* feat: ci

* feat: fix novita models prices
2026-01-14 03:30:19 +05:30
Yuta Saito
22aad95bb1 fix: rest_endpoints own allowed_mcp_servers for MCP calls 2026-01-14 06:39:46 +09:00
Sameer Kankute
93203cda7c
Merge pull request #19027 from BerriAI/litellm_add_0_budget_model_bypass
[Feat] Add support for 0 cost models
2026-01-13 18:05:36 +05:30
Sameer Kankute
e98c2e4425
Merge pull request #19012 from BerriAI/litellm_fix_model_deployment_routing
Fix: Model matching priority in configuration
2026-01-13 17:55:01 +05:30
Sameer Kankute
bb0ab38636
Merge pull request #19007 from BerriAI/litellm_fix_header_forwarding_passthrough
Fix: Header forwarding in bedrock passthrough
2026-01-13 17:50:14 +05:30
Sameer Kankute
54f6f55c98
Merge pull request #18873 from BerriAI/litellm_staging_01_09_2026
staging 01/09/2025
2026-01-13 17:31:44 +05:30
Sameer Kankute
f2cb861d6a
Merge pull request #18976 from BerriAI/litellm_staging_12_19_2025
Staging 12/19/2025 - implement failopen option default to True on grayswan guardrail (#18266)
2026-01-13 17:00:36 +05:30
Sameer Kankute
de6330b6b6 Fix test_async_otel_callback[False] 2026-01-13 16:59:17 +05:30
Sameer Kankute
1932d03aed Add docs on Zero-Cost Models 2026-01-13 16:44:02 +05:30
Sameer Kankute
762a3ef090 Add support for 0 cost models 2026-01-13 16:39:57 +05:30
Sameer Kankute
d656f01bc9
Merge pull request #19009 from Dima-Mediator/fix-image-tokens-spend-logging
Fix image tokens spend logging for /images/generations
2026-01-13 15:03:37 +05:30
Sameer Kankute
5349d8922c
Merge pull request #19003 from BerriAI/litellm_add_azure_ai_claude_opus
Add pricing of azure_ai/claude-opus-4-5
2026-01-13 13:52:47 +05:30
Yuta Saito
658bbcc2d5 fix: mcp rest auth check 2026-01-13 17:18:58 +09:00
YutaSaito
b3e126222f
Merge pull request #19013 from BerriAI/litellm_test_comment_out_flaky
[test] temporarily disable flaky responses_id_security tests
2026-01-13 15:52:57 +09:00
Yuta Saito
2c8ac2c3f1 test: temporarily disable flaky responses_id_security tests 2026-01-13 15:51:37 +09:00
Sameer Kankute
ecb3959c3c
Merge pull request #18208 from Chesars/fix/case-insensitive-model-cost-lookup
fix: case-insensitive model cost map lookup
2026-01-13 11:53:43 +05:30
Sameer Kankute
dfece51f8c Fix: Model matching priority in configuration 2026-01-13 11:44:48 +05:30
yuneng-jiang
da902c5c55 Migrate User and Team filters to use reusable components 2026-01-12 21:04:16 -08:00
Sameer Kankute
005541075b Fix: Header forwarding in bedrock passthrough 2026-01-13 09:45:14 +05:30
Dima-Mediator
7c61933bc5 Fix image tokens spend logging for /images/generations 2026-01-12 23:07:08 -05:00
Sameer Kankute
5a51b74658 Add pricing of azure_ai/claude-opus-4-5 2026-01-13 09:15:05 +05:30