BYOK is live — 1,000,000 free BYOK requests every month, no top-up required Learn more →

Model catalog

Every model below is served through the same OpenAI-compatible endpoint with one API key. Prices are in USD per 1M tokens; the console shows the live price applied to each request. The catalog grows as new providers are onboarded.

Anthropic (Claude)12 models

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
claude-fable-5$9.5$47.5$0.95
claude-opus-4-8-fast$9.5$47.5$0.95
claude-opus-5-fast$9.5$47.5$0.95
claude-opus-4.8$5.7$28.5$0.57
claude-opus-4.5$4.75$23.75$0.475
claude-opus-4.6$4.75$23.75$0.475
claude-opus-4.7$4.75$23.75$0.475
claude-opus-5 details$4.75$23.75$0.475
claude-sonnet-4.5$2.85$14.25$0.285
claude-sonnet-4.6$2.85$14.25$0.285
claude-sonnet-5 details$1.9$9.5$0.19
claude-haiku-4.5 details$0.95$4.75$0.095

OpenAI (GPT)6 models

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
gpt-5.5$4.75$28.5$0.475
gpt-5.6-sol$4.75$28.5$0.475
gpt-5.4 details$2.375$14.25$0.2375
gpt-5.6-terra$2.375$14.25$0.2375
gpt-5.6-luna$0.95$5.7$0.095
gpt-5.4-mini$0.7125$4.275$0.07125

Google (Gemini)8 models

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
gemini-3.1-pro$2$12$0.2
gemini-3.1-pro-preview$1.9$11.4$0.19
nano-banana-pro$1.9$11.4$0.19
gemini-3-5-flash$1.425$8.55$0.1425
gemini-3-flash-preview$0.475$2.85$0.0475
nano-banana-2$0.475$2.85$0.0475
gemini-2.5-flash$0.285$2.375$0.0285
gemini-3.1-flash-lite$0.2375$1.425$0.02375

DeepSeek2 models

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
deepseek-v4-pro details$0.41325$0.8265$0.041325
deepseek-v4-flash details$0.133$0.266$0.0133

GLM (Zhipu)4 models

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
glm-5.2 details$1.045$3.284286$0.1045
glm-5$0.9025$2.4225$0.09025
glm-4.5$0.57$2.09$0.057
glm-4.6$0.475$1.9$0.0475

xAI (Grok)1 model

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
grok-4.5$1.9$5.7$0.19

Kimi (Moonshot)3 models

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
kimi-k2.6$0.9025$3.8$0.152
kimi-k2.7-code details$0.9025$3.8$0.1805
kimi-k2.5$0.57$2.85$0.095

MiniMax1 model

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
minimax-m2.7 details$0.228$0.912$0.0228

Qwen (Alibaba)8 models

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
qwen3.7-max$2.5$7.5$0.5
qwen3-coder-plus details ≤120K$1.8$9$0.36
qwen3.5-397b$0.6$3.6$0.12
qwen3-coder-flash ≤120K$0.5$2.5$0.1
qwen3.5-plus$0.5$3$0.1
qwen3.7-plus ≤240K$0.4$1.6$0.08
qwen3.5-flash$0.1$0.4$0.02
qwen3.7-flash ≤240K$0.1$0.4$0.02

aion1 model

ModelInput · $ / 1M tokensOutput · $ / 1M tokensCache read · $ / 1M tokens
aion-2.0$0.8$1.6$0.08

Prices as currently published; the live per-request price in the console prevails. Cached input tokens are billed at the cache-read price (10% of input); cache writes carry no separate charge — they are billed as normal input.

Context caps: models marked ≤120K / ≤240K accept inputs up to that size, at a single flat rate matching Alibaba Cloud's official price for the corresponding usage tier. Larger inputs return a context_length_exceeded error before any charge. For 1M-context work, qwen3.7-max carries the official single rate with no cap.

Models with a details link have their own page — price, a ready-to-run call, and the closest alternatives. Those pages are rolling out model by model; every model in the table is callable today either way.

Cost calculator

Get your API key See live pricing in the console

Last updated: 2026-07-29