jishie

Llama 3.3 70b Instruct

nvidia/meta/llama-3.3-70b-instruct · in: text → out: text · price provider-declared (via models.dev). Capabilities: tools text.

free
Blended $/1M (3:1)
free
Input $/1M
free
Output $/1M
free
1M in + 1M out
—
Cache read $/1M
#317 of 7,708
Cheapest rank
Provider
Nvidia (nvidia)
Kind
text (outputs text)
Context
128,000 tokens
Max output
4,096 tokens
Capabilities
tools text

Same model, other providers

“Llama 3.3 70b Instruct” is offered by 10 providers — cheapest-effective first. This one (Nvidia) is the cheapest.

ProviderInputOutputBlendedContext
Nvidia (this)freefree$0.00128k
NanoGPT$0.05$0.23$0.10131k
Meganova$0.1$0.3$0.15131k
IO.NET$0.13$0.38$0.19128k
NovitaAI$0.135$0.4$0.20131k
Merge Gateway$0.22$0.5$0.29131k
Amazon Bedrock$0.72$0.72$0.72128k
Vertex$0.72$0.72$0.72128k
Regolo AI$0.6$2.7$1.13128k
CloudFerro Sherlock$2.92$2.92$2.9270k

Call it

jishie indexes & prices models; it does not proxy inference.

Most providers are OpenAI-compatible — point base_url at Nvidia and pass this model id:

curl $BASE_URL/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"meta/llama-3.3-70b-instruct","messages":[{"role":"user","content":"hello"}]}'

Route programmatically with MCP find_model / get_model, or fetch this record free (no wallet):

← all Nvidia models · model index · JSON