Eddie Richter
900a57f5b1
Adding new models to the lemonade provider
2025-10-14 16:56:36 -06:00
Ishaan Jaff
a10425d1ce
[Feat] Allow Team Admins to export a report of the team spending ( #15542 )
...
* v0 for export
* v0 for Export
* add types for EntityUsageExportModalProps
* add folder struct
* add summry selector
* add selector for export
* add utils for entity usage export
* refactored buttons
* fixes name
* fix alignment
* test: EntityUsageExportModal
* fix lint
2025-10-14 15:07:59 -07:00
Ishaan Jaff
65163c7ccb
[Fix] GEMINI - CLI - add google_routes to llm_api_routes ( #15500 )
...
* fix: add google_routes to llm_api_routes
* test: test_virtual_key_llm_api_routes_allows_google_routes
2025-10-14 13:57:39 -07:00
AlexsanderHamir
5bffa58ee0
feat(ssl): add configurable ECDH curve for TLS performance
...
Add `ssl_ecdh_curve` setting to configure TLS key exchange curve.
Allows disabling PQC on OpenSSL 3.x for better performance.
Configurable via SDK (litellm.ssl_ecdh_curve), YAML (litellm_settings),
or env var (SSL_ECDH_CURVE). Common curves: X25519, prime256v1, secp384r1.
2025-10-14 13:57:39 -07:00
Krrish Dholakia
b0d963cc0c
docs(index.md): bump rc
2025-10-14 13:57:39 -07:00
AlexsanderHamir
fcb85f8856
perf(router): optimize timing functions in completion hot path
...
Replace time.time() with more appropriate timing functions for better
performance and reliability:
- Use time.perf_counter() for duration measurements in acompletion(),
_acompletion(), and async_get_available_deployment()
- Use time.monotonic() for timeout calculations in scheduler methods
(schedule_acompletion and _schedule_factory)
Benefits:
- 30-40% faster timing calls (~300ns savings per call)
- time.monotonic() provides reliable timeouts unaffected by system
clock changes (NTP adjustments, DST, manual time changes)
- time.perf_counter() offers highest resolution for performance metrics
- Follows Python best practices for timing operations
2025-10-14 13:57:39 -07:00
Krrish Dholakia
cb29e33cad
docs: fix doc
2025-10-14 13:57:39 -07:00
Dhruv Yadav
b57406e53e
add tests for openrouter cost tracking
2025-10-14 13:57:39 -07:00
Dhruv Yadav
09662b5081
direct cost calculation from openrouter
2025-10-14 13:57:39 -07:00
AlexsanderHamir
2e57d19a55
update benchmarks
2025-10-14 13:57:39 -07:00
AlexsanderHamir
6df051e87b
docs: update benchmark results with improved infrastructure
...
- Update to 2 instance baseline (1035 RPS @ 200ms median)
- Add 4 instance results (1170 RPS @ 100ms median)
- Update machine specs to 4 CPU / 8GB RAM
- Update Locust settings to 1000 users
2025-10-14 13:57:39 -07:00
kowyo
92bd238b87
fix comment
2025-10-14 13:57:39 -07:00
kowyo
6896c60b5b
fix: only use think level for gpt-oss model
2025-10-14 13:57:39 -07:00
Kowyo
9557a37d37
fix: update 'think' parameter assignment to use provided value in transformation.py
2025-10-14 13:57:39 -07:00
Kowyo
e30b716ba8
fix: add 'think' parameter handling in ollama_chat.py
2025-10-14 13:57:39 -07:00
Kowyo
4f33bc4e46
fix(ollama/chat): 'think' param handling
2025-10-14 13:57:39 -07:00
mubashir1osmani
6424480571
add ecs to docs
2025-10-14 13:57:39 -07:00
huangyafei
9a980f36d4
Add anthropic/claude-sonnet-4.5 to OpenRouter cost map
2025-10-14 13:57:39 -07:00
Davi S. Zucon
9d944b1404
small fix code snippet custom_prompt_management.md
...
passing parameter:
prompt_id directly to openai client, raises error:
TypeError: Completions.create() got an unexpected keyword argument 'prompt_id'
instead use:
extra_body={
"prompt_id": "1234"
}
2025-10-14 16:37:00 -03:00
Ishaan Jaffer
c86c6bc507
v0 for Export
2025-10-14 11:11:37 -07:00
Ishaan Jaffer
bda58bc3ed
v0 for export
2025-10-14 11:11:24 -07:00
Felipe Gare
0212eb6f04
change gpt-5-codex support in model_price json
2025-10-14 13:37:56 -03:00
Hana Volků
55e44e1b1d
Prompt caching for anthropic models with openrouter
2025-10-14 14:51:55 +02:00
Krrish Dholakia
3e193a3542
bump: version 1.78.0 → 1.78.1
2025-10-13 14:26:47 -07:00
Krrish Dholakia
3d7c55516e
build: bump version
2025-10-13 14:26:38 -07:00
Ishaan Jaff
f13eb283e1
[Fix] GEMINI - CLI - add google_routes to llm_api_routes ( #15500 )
...
* fix: add google_routes to llm_api_routes
* test: test_virtual_key_llm_api_routes_allows_google_routes
2025-10-13 10:58:47 -07:00
Krrish Dholakia
0c86c3d791
docs(index.md): bump rc
2025-10-13 10:56:14 -07:00
Krrish Dholakia
611bda94cb
docs: fix doc
2025-10-13 09:59:07 -07:00
Classic298
732063074d
correct claude opus
2025-10-13 14:20:24 +02:00
Krish Dholakia
2ea7005c40
Merge pull request #15448 from dhruvyad/main
...
Get completion cost directly from OpenRouter
2025-10-12 22:12:52 -07:00
Krish Dholakia
d81843070a
Merge pull request #15461 from BerriAI/litellm_update_docs
...
[Docs] - Update benchmark results
2025-10-12 22:11:57 -07:00
Krish Dholakia
760000e48e
Merge pull request #15465 from kowyo/kowyo/fix-ollama-think
...
fix(ollama/chat): correctly map reasoning_effort to think in requests
2025-10-12 22:11:17 -07:00
Krish Dholakia
b7b16d69bc
Merge pull request #15468 from mubashir1osmani/litellm_add_docs
...
docs: add ecs deployment guide
2025-10-12 22:10:00 -07:00
Krish Dholakia
176c45d51b
Merge pull request #15472 from huangyafei/update_price
...
Add anthropic/claude-sonnet-4.5 to OpenRouter cost map
2025-10-12 22:07:04 -07:00
Krish Dholakia
ff20f8402a
Merge pull request #15456 from BerriAI/litellm_staging_branch_10_11_2025_p1
...
Litellm staging branch 10 11 2025 p1
2025-10-12 22:01:45 -07:00
Krrish Dholakia
0ffc81f010
fix: fix reformatting
2025-10-12 22:01:16 -07:00
Krrish Dholakia
1949436047
test: update test
2025-10-12 21:58:06 -07:00
Krish Dholakia
04dc9d091c
Merge branch 'main' into litellm_staging_branch_10_11_2025_p1
2025-10-12 21:55:44 -07:00
huangyafei
409299de84
Add anthropic/claude-sonnet-4.5 to OpenRouter cost map
2025-10-13 10:46:37 +08:00
mubashir1osmani
950ee924fc
add ecs to docs
2025-10-12 16:13:56 -04:00
Dhruv Yadav
84a65440c5
add tests for openrouter cost tracking
2025-10-12 12:36:24 +05:30
Kowyo
89aad0e21f
Merge branch 'BerriAI:main' into kowyo/fix-ollama-think
2025-10-12 11:54:30 +08:00
Ishaan Jaffer
e761665f22
docs fix
2025-10-11 18:21:35 -07:00
AlexsanderHamir
5d1fa76d14
update benchmarks
2025-10-11 18:06:37 -07:00
AlexsanderHamir
2cb1db27cb
docs: update benchmark results with improved infrastructure
...
- Update to 2 instance baseline (1035 RPS @ 200ms median)
- Add 4 instance results (1170 RPS @ 100ms median)
- Update machine specs to 4 CPU / 8GB RAM
- Update Locust settings to 1000 users
2025-10-11 18:00:26 -07:00
Imad Saddik
133189e191
Fixed a few typos ( #15267 )
2025-10-11 17:53:42 -07:00
Hampus Näsström
e650f39821
Reduce claude-4-sonnet max_output_tokens to 64k ( #15409 )
...
Claude 4 Sonnet doesn't support more than 64k output tokens: https://cloud.google.com/vertex-ai/generative-ai/docs/partner-models/claude/sonnet-4
2025-10-11 17:51:27 -07:00
Copilot
f5359ba007
Fix apply_guardrail endpoint returning raw string instead of ApplyGuardrailResponse ( #15436 )
...
* Initial plan
* Fix apply_guardrail endpoint to return ApplyGuardrailResponse
Co-authored-by: ishaan-jaff <29436595+ishaan-jaff@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: ishaan-jaff <29436595+ishaan-jaff@users.noreply.github.com>
2025-10-11 17:50:37 -07:00
Alexsander Hamir
300a142926
[Add]: perf summary ( #15458 )
...
* add: perf summary
* update docs
* fix: focus on p99
2025-10-11 17:38:13 -07:00
Krrish Dholakia
78e2274381
fix(pass_through_endpoints.py): use path instead of endpoint
...
includes the mapped route
2025-10-11 17:04:56 -07:00