jishie

GLM-5.3-Flash

pareto/z-ai/glm-5.3-flash · family glm-flash · in: text, image, video → out: text · price provider-declared (via models.dev). Capabilities: tools reasoning text image video.

$0.05
Blended $/1M (3:1)
$0.03
Input $/1M
$0.1
Output $/1M
$0.13
1M in + 1M out
$0.006 −80%
Cache read $/1M
#685 of 7,795
Cheapest rank
Provider
Pareto Inference (pareto)
Kind
text (outputs text)
Context
1,000,000 tokens
Max output
131,072 tokens
Capabilities
tools reasoning text image video

Same model, other providers

“GLM-5.3-Flash” is offered by 48 providers — cheapest-effective first. NaN is the cheapest at $0.00 blended.

ProviderInputOutputBlendedContext
NaNfreefree$0.001M
Zhipu AI Coding Planfreefree$0.001M
Volcengine Ark Coding Planfreefree$0.001M
Z.AI Coding Planfreefree$0.001M
SCNet Token Planfreefree$0.001M
Nvidiafreefree$0.001M
Kenarifreefree$0.001M
Pareto Inference (this)$0.03$0.1$0.051M
TokenGo$0.075$0.025$0.061M
302.AI$0.075$0.25$0.121M
OrcaRouter$0.075$0.25$0.121M
DevPass (LLM Gateway)$0.088$0.25$0.131M
Cortecs$0.1$0.35$0.161M

Call it

jishie indexes & prices models; it does not proxy inference.

Most providers are OpenAI-compatible — point base_url at Pareto Inference and pass this model id:

curl $BASE_URL/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"z-ai/glm-5.3-flash","messages":[{"role":"user","content":"hello"}]}'

Route programmatically with MCP find_model / get_model, or fetch this record free (no wallet):

← all Pareto Inference models · model index · JSON