jishie

Nemotron 3.5 Lightning 30B A3B

runinfra/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 · family nemotron · in: text → out: text · price provider-declared (via models.dev). Capabilities: tools reasoning text.

$0.08
Blended $/1M (3:1)
$0.05
Input $/1M
$0.15
Output $/1M
$0.20
1M in + 1M out
$0.01 −80%
Cache read $/1M
#834 of 7,876
Cheapest rank
Provider
RunInfra (runinfra)
Kind
text (outputs text)
Context
262,144 tokens
Max output
32,768 tokens
Capabilities
tools reasoning text

Same model, other providers

“Nemotron 3.5 Lightning 30B A3B” is offered by 8 providers — cheapest-effective first. Merge Gateway is the cheapest at $0.00 blended.

ProviderInputOutputBlendedContext
Merge Gatewayfreefree$0.001M
Nvidiafreefree$0.00262k
Kilo Gateway$0.039$0.18$0.07262k
RunInfra (this)$0.05$0.15$0.08262k
OpenRouter$0.0595$0.17$0.09262k
Fireworks AI$0.05$0.2$0.09262k
Nebius Token Factory$0.06$0.24$0.101M
Pioneer$0.5$0.5$0.508k

Call it

jishie indexes & prices models; it does not proxy inference.

Most providers are OpenAI-compatible — point base_url at RunInfra and pass this model id:

curl $BASE_URL/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16","messages":[{"role":"user","content":"hello"}]}'

Route programmatically with MCP find_model / get_model, or fetch this record free (no wallet):

← all RunInfra models · model index · JSON