jishie

Llama 3.3 Nemotron Super 49B v1.5

deepinfra/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5 · family nemotron · in: text → out: text · price provider-declared (via models.dev). Capabilities: tools reasoning text.

$0.40
Blended $/1M (3:1)
$0.4
Input $/1M
$0.4
Output $/1M
$0.80
1M in + 1M out
—
Cache read $/1M
#2462 of 7,713
Cheapest rank
Provider
Deep Infra (deepinfra)
Kind
text (outputs text)
Context
131,072 tokens
Max output
131,072 tokens
Capabilities
tools reasoning text

Same model, other providers

“Llama 3.3 Nemotron Super 49B v1.5” is offered by 2 providers — cheapest-effective first. Nvidia is the cheapest at $0.00 blended.

ProviderInputOutputBlendedContext
Nvidiafreefree$0.00131k
Deep Infra (this)$0.4$0.4$0.40131k

Call it

jishie indexes & prices models; it does not proxy inference.

Most providers are OpenAI-compatible — point base_url at Deep Infra and pass this model id:

curl $BASE_URL/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"nvidia/Llama-3.3-Nemotron-Super-49B-v1.5","messages":[{"role":"user","content":"hello"}]}'

Route programmatically with MCP find_model / get_model, or fetch this record free (no wallet):

← all Deep Infra models · model index · JSON