Commit graph

26846 commits

Author SHA1 Message Date
Eddie Richter
900a57f5b1 Adding new models to the lemonade provider 2025-10-14 16:56:36 -06:00
Ishaan Jaff
a10425d1ce
[Feat] Allow Team Admins to export a report of the team spending (#15542)
* v0 for export

* v0 for Export

* add types for EntityUsageExportModalProps

* add folder struct

* add summry selector

* add selector for export

* add utils for entity usage export

* refactored buttons

* fixes name

* fix alignment

* test: EntityUsageExportModal

* fix lint
2025-10-14 15:07:59 -07:00
Ishaan Jaff
65163c7ccb [Fix] GEMINI - CLI - add google_routes to llm_api_routes (#15500)
* fix: add google_routes to llm_api_routes

* test: test_virtual_key_llm_api_routes_allows_google_routes
2025-10-14 13:57:39 -07:00
AlexsanderHamir
5bffa58ee0 feat(ssl): add configurable ECDH curve for TLS performance
Add `ssl_ecdh_curve` setting to configure TLS key exchange curve.
Allows disabling PQC on OpenSSL 3.x for better performance.

Configurable via SDK (litellm.ssl_ecdh_curve), YAML (litellm_settings),
or env var (SSL_ECDH_CURVE). Common curves: X25519, prime256v1, secp384r1.
2025-10-14 13:57:39 -07:00
Krrish Dholakia
b0d963cc0c docs(index.md): bump rc 2025-10-14 13:57:39 -07:00
AlexsanderHamir
fcb85f8856 perf(router): optimize timing functions in completion hot path
Replace time.time() with more appropriate timing functions for better
performance and reliability:

- Use time.perf_counter() for duration measurements in acompletion(),
  _acompletion(), and async_get_available_deployment()
- Use time.monotonic() for timeout calculations in scheduler methods
  (schedule_acompletion and _schedule_factory)

Benefits:
- 30-40% faster timing calls (~300ns savings per call)
- time.monotonic() provides reliable timeouts unaffected by system
  clock changes (NTP adjustments, DST, manual time changes)
- time.perf_counter() offers highest resolution for performance metrics
- Follows Python best practices for timing operations
2025-10-14 13:57:39 -07:00
Krrish Dholakia
cb29e33cad docs: fix doc 2025-10-14 13:57:39 -07:00
Dhruv Yadav
b57406e53e add tests for openrouter cost tracking 2025-10-14 13:57:39 -07:00
Dhruv Yadav
09662b5081 direct cost calculation from openrouter 2025-10-14 13:57:39 -07:00
AlexsanderHamir
2e57d19a55 update benchmarks 2025-10-14 13:57:39 -07:00
AlexsanderHamir
6df051e87b docs: update benchmark results with improved infrastructure
- Update to 2 instance baseline (1035 RPS @ 200ms median)
- Add 4 instance results (1170 RPS @ 100ms median)
- Update machine specs to 4 CPU / 8GB RAM
- Update Locust settings to 1000 users
2025-10-14 13:57:39 -07:00
kowyo
92bd238b87 fix comment 2025-10-14 13:57:39 -07:00
kowyo
6896c60b5b fix: only use think level for gpt-oss model 2025-10-14 13:57:39 -07:00
Kowyo
9557a37d37 fix: update 'think' parameter assignment to use provided value in transformation.py 2025-10-14 13:57:39 -07:00
Kowyo
e30b716ba8 fix: add 'think' parameter handling in ollama_chat.py 2025-10-14 13:57:39 -07:00
Kowyo
4f33bc4e46 fix(ollama/chat): 'think' param handling 2025-10-14 13:57:39 -07:00
mubashir1osmani
6424480571 add ecs to docs 2025-10-14 13:57:39 -07:00
huangyafei
9a980f36d4 Add anthropic/claude-sonnet-4.5 to OpenRouter cost map 2025-10-14 13:57:39 -07:00
Davi S. Zucon
9d944b1404
small fix code snippet custom_prompt_management.md
passing parameter: 
prompt_id directly to openai client, raises error: 
TypeError: Completions.create() got an unexpected keyword argument 'prompt_id'

instead use: 
extra_body={
        "prompt_id": "1234"
 }
2025-10-14 16:37:00 -03:00
Ishaan Jaffer
c86c6bc507 v0 for Export 2025-10-14 11:11:37 -07:00
Ishaan Jaffer
bda58bc3ed v0 for export 2025-10-14 11:11:24 -07:00
Felipe Gare
0212eb6f04 change gpt-5-codex support in model_price json 2025-10-14 13:37:56 -03:00
Hana Volků
55e44e1b1d
Prompt caching for anthropic models with openrouter 2025-10-14 14:51:55 +02:00
Krrish Dholakia
3e193a3542 bump: version 1.78.0 → 1.78.1 2025-10-13 14:26:47 -07:00
Krrish Dholakia
3d7c55516e build: bump version 2025-10-13 14:26:38 -07:00
Ishaan Jaff
f13eb283e1
[Fix] GEMINI - CLI - add google_routes to llm_api_routes (#15500)
* fix: add google_routes to llm_api_routes

* test: test_virtual_key_llm_api_routes_allows_google_routes
2025-10-13 10:58:47 -07:00
Krrish Dholakia
0c86c3d791 docs(index.md): bump rc 2025-10-13 10:56:14 -07:00
Krrish Dholakia
611bda94cb docs: fix doc 2025-10-13 09:59:07 -07:00
Classic298
732063074d
correct claude opus 2025-10-13 14:20:24 +02:00
Krish Dholakia
2ea7005c40
Merge pull request #15448 from dhruvyad/main
Get completion cost directly from OpenRouter
2025-10-12 22:12:52 -07:00
Krish Dholakia
d81843070a
Merge pull request #15461 from BerriAI/litellm_update_docs
[Docs] - Update benchmark results
2025-10-12 22:11:57 -07:00
Krish Dholakia
760000e48e
Merge pull request #15465 from kowyo/kowyo/fix-ollama-think
fix(ollama/chat): correctly map reasoning_effort to think in requests
2025-10-12 22:11:17 -07:00
Krish Dholakia
b7b16d69bc
Merge pull request #15468 from mubashir1osmani/litellm_add_docs
docs: add ecs deployment guide
2025-10-12 22:10:00 -07:00
Krish Dholakia
176c45d51b
Merge pull request #15472 from huangyafei/update_price
Add anthropic/claude-sonnet-4.5 to OpenRouter cost map
2025-10-12 22:07:04 -07:00
Krish Dholakia
ff20f8402a
Merge pull request #15456 from BerriAI/litellm_staging_branch_10_11_2025_p1
Litellm staging branch 10 11 2025 p1
2025-10-12 22:01:45 -07:00
Krrish Dholakia
0ffc81f010 fix: fix reformatting 2025-10-12 22:01:16 -07:00
Krrish Dholakia
1949436047 test: update test 2025-10-12 21:58:06 -07:00
Krish Dholakia
04dc9d091c
Merge branch 'main' into litellm_staging_branch_10_11_2025_p1 2025-10-12 21:55:44 -07:00
huangyafei
409299de84 Add anthropic/claude-sonnet-4.5 to OpenRouter cost map 2025-10-13 10:46:37 +08:00
mubashir1osmani
950ee924fc add ecs to docs 2025-10-12 16:13:56 -04:00
Dhruv Yadav
84a65440c5 add tests for openrouter cost tracking 2025-10-12 12:36:24 +05:30
Kowyo
89aad0e21f
Merge branch 'BerriAI:main' into kowyo/fix-ollama-think 2025-10-12 11:54:30 +08:00
Ishaan Jaffer
e761665f22 docs fix 2025-10-11 18:21:35 -07:00
AlexsanderHamir
5d1fa76d14 update benchmarks 2025-10-11 18:06:37 -07:00
AlexsanderHamir
2cb1db27cb docs: update benchmark results with improved infrastructure
- Update to 2 instance baseline (1035 RPS @ 200ms median)
- Add 4 instance results (1170 RPS @ 100ms median)
- Update machine specs to 4 CPU / 8GB RAM
- Update Locust settings to 1000 users
2025-10-11 18:00:26 -07:00
Imad Saddik
133189e191
Fixed a few typos (#15267) 2025-10-11 17:53:42 -07:00
Hampus Näsström
e650f39821
Reduce claude-4-sonnet max_output_tokens to 64k (#15409)
Claude 4 Sonnet doesn't support more than 64k output tokens: https://cloud.google.com/vertex-ai/generative-ai/docs/partner-models/claude/sonnet-4
2025-10-11 17:51:27 -07:00
Copilot
f5359ba007
Fix apply_guardrail endpoint returning raw string instead of ApplyGuardrailResponse (#15436)
* Initial plan

* Fix apply_guardrail endpoint to return ApplyGuardrailResponse

Co-authored-by: ishaan-jaff <29436595+ishaan-jaff@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: ishaan-jaff <29436595+ishaan-jaff@users.noreply.github.com>
2025-10-11 17:50:37 -07:00
Alexsander Hamir
300a142926
[Add]: perf summary (#15458)
* add: perf summary

* update docs

* fix: focus on p99
2025-10-11 17:38:13 -07:00
Krrish Dholakia
78e2274381 fix(pass_through_endpoints.py): use path instead of endpoint
includes the mapped route
2025-10-11 17:04:56 -07:00