feat(model_prices): add NVIDIA Nemotron 3.5 Lightning on OpenRouter and DeepInfra (#36696)

* feat(model_prices): add NVIDIA Nemotron 3.5 Lightning on OpenRouter and DeepInfra

Nemotron 3.5 Lightning shipped 2026-08-11 with public per-token pricing on
OpenRouter and DeepInfra at $0.05/M in and $0.20/M out. Without cost map
entries both ids raise "This model isn't mapped yet" and log at zero spend.

* fix(model_prices): stop asserting an output cap for Nemotron 3.5 Lightning

262144 is the native context window, not the output budget, and neither
OpenRouter nor DeepInfra publishes an output cap. Keeps max_input_tokens at the
native 256K window: 1M needs VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 on a self-hosted
deployment, so it is not what these hosted endpoints serve.

* chore(tests): drop the Nemotron 3.5 Lightning metadata test

Requested on the review thread: the cost map entries stand on their own.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
This commit is contained in:
devin-ai-integration[bot] 2026-08-12 21:48:02 +00:00 • committed by GitHub
parent 1911269ddf
commit e52ae039f2
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
2 changed files with 44 additions and 0 deletions

View file

@ -15552,6 +15552,17 @@
"supports_tool_choice": true,
"supports_function_calling": true
},
"deepinfra/nvidia/NVIDIA-Nemotron-3.5-Lightning": {
"max_input_tokens": 262144,
"input_cost_per_token": 5e-08,
"output_cost_per_token": 2e-07,
"litellm_provider": "deepinfra",
"mode": "chat",
"source": "https://deepinfra.com/nvidia/NVIDIA-Nemotron-3.5-Lightning",
"supports_tool_choice": true,
"supports_function_calling": true,
"supports_reasoning": true
},
"deepinfra/nvidia/NVIDIA-Nemotron-Nano-9B-v2": {
"max_tokens": 131072,
"max_input_tokens": 131072,
@ -31754,6 +31765,17 @@
"supports_video_input": true,
"supports_vision": true
},
"openrouter/nvidia/nemotron-3.5-lightning": {
"input_cost_per_token": 5e-08,
"litellm_provider": "openrouter",
"max_input_tokens": 262144,
"mode": "chat",
"output_cost_per_token": 2e-07,
"source": "https://openrouter.ai/nvidia/nemotron-3.5-lightning",
"supports_function_calling": true,
"supports_reasoning": true,
"supports_tool_choice": true
},
"openrouter/openai/gpt-3.5-turbo": {
"input_cost_per_token": 1.5e-06,
"litellm_provider": "openrouter",

View file

@ -15552,6 +15552,17 @@
"supports_tool_choice": true,
"supports_function_calling": true
},
"deepinfra/nvidia/NVIDIA-Nemotron-3.5-Lightning": {
"max_input_tokens": 262144,
"input_cost_per_token": 5e-08,
"output_cost_per_token": 2e-07,
"litellm_provider": "deepinfra",
"mode": "chat",
"source": "https://deepinfra.com/nvidia/NVIDIA-Nemotron-3.5-Lightning",
"supports_tool_choice": true,
"supports_function_calling": true,
"supports_reasoning": true
},
"deepinfra/nvidia/NVIDIA-Nemotron-Nano-9B-v2": {
"max_tokens": 131072,
"max_input_tokens": 131072,
@ -31754,6 +31765,17 @@
"supports_video_input": true,
"supports_vision": true
},
"openrouter/nvidia/nemotron-3.5-lightning": {
"input_cost_per_token": 5e-08,
"litellm_provider": "openrouter",
"max_input_tokens": 262144,
"mode": "chat",
"output_cost_per_token": 2e-07,
"source": "https://openrouter.ai/nvidia/nemotron-3.5-lightning",
"supports_function_calling": true,
"supports_reasoning": true,
"supports_tool_choice": true
},
"openrouter/openai/gpt-3.5-turbo": {
"input_cost_per_token": 1.5e-06,
"litellm_provider": "openrouter",