Cached at:
08/19/26, 03:34 PM
# GLM-5.3 hits the API at $1.4/$4.4 per million tokens
Source: [https://venturebeat.com/ai/glm-5-3-hits-the-api-at-1-4-4-4-per-million-tokens](https://venturebeat.com/ai/glm-5-3-hits-the-api-at-1-4-4-4-per-million-tokens)
After a[stunning debut last week](https://venturebeat.com/technology/glm-5-3-is-here-with-advanced-cyber-capabilities-and-reportedly-already-found-a-serious-vulnerability-in-cursor)with cyber capabilities so advanced they reportedly found a previously undetected vulnerability in Cursor, GLM\-5\.3, the new frontier open source language model from Chinese startup z\.ai, has[now hit the application programming interface \(API\)](https://x.com/Zai_org/status/2089816129011098048?s=20)— allowing developers the ability to build atop it and plug it into their agents and applications\.
Developers who previously subscribed to a GLM Coding Plan are currently limited to the OpenAI Chat Completions\-compatible protocol\. Z\.ai said it plans to make the model's weights openly available, but a precise date and licensing remain to be seen\.
On the API, the price is unchanged from**GLM\-5\.2: $1\.40 per million input tokens and $4\.40 per million output tokens**\. Cached input costs $0\.26 per million tokens, while Z\.ai currently lists cached\-input storage as free for a limited time\.
That means developers can move to the new generation without taking a higher posted per\-token rate from Z\.ai, even as the company claims substantially stronger coding and long\-horizon agent performance\. At those rates, GLM\-5\.3 sits well below several of the highest\-end frontier APIs\.
**Model**
**Input \($/1M\)**
**Output \($/1M\)**
**Total \($/1M\)**
**Source**
Muse Spark 1\.2 Contributor
$0\.10
$0\.20
$0\.30
[Meta](https://dev.meta.ai/docs/pricing-rate-limits)
MiMo\-V2\.5 Flash
$0\.10
$0\.30
$0\.40
[Xiaomi](https://platform.xiaomimimo.com/docs/en-US/pricing)
DeepSeek\-V4\-Flash — off\-peak
$0\.22
$0\.66
$0\.88
[DeepSeek](https://x.com/deepseek_ai/status/2087864589895798968)
GPT\-5\.6 Luna
$0\.20
$1\.20
$1\.40
[OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)
MiniMax\-M3
$0\.30
$1\.20
$1\.50
[MiniMax](https://platform.minimax.io/subscribe/token-plan?tab=api-enterprise)
LongCat\-2\.0 — limited\-time promo
$0\.30
$1\.20
$1\.50
[LongCat](https://longcat.chat/platform/docs/APIPayAsYouGo.html)
DeepSeek\-V4\-Flash — peak hours
$0\.44
$1\.32
$1\.76
[DeepSeek](https://x.com/deepseek_ai/status/2087864589895798968)
MiMo\-V2\.5
$0\.40
$2\.00
$2\.40
[Xiaomi](https://platform.xiaomimimo.com/docs/en-US/pricing)
DeepSeek\-V4\-Pro — off\-peak
$0\.66
$1\.98
$2\.64
[DeepSeek](https://x.com/deepseek_ai/status/2087864589895798968)
LongCat\-2\.0 — standard
$0\.75
$2\.95
$3\.70
[LongCat](https://longcat.chat/platform/docs/APIPayAsYouGo.html)
MiMo\-V2\.5 Pro \(≤256K\)
$1\.00
$3\.00
$4\.00
[Xiaomi](https://platform.xiaomimimo.com/docs/en-US/pricing)
Gemini 3\.6 Flash — through Dec\. 31, 2026
$0\.75
$3\.75
$4\.50
[Google](https://ai.google.dev/gemini-api/docs/pricing)
Gemini 3\.7 Flash — through Dec\. 31, 2026
$0\.75
$3\.75
$4\.50
[Google](https://ai.google.dev/gemini-api/docs/pricing)
DeepSeek\-V4\-Pro — peak hours
$1\.32
$3\.96
$5\.28
[DeepSeek](https://x.com/deepseek_ai/status/2087864589895798968)
Muse Spark 1\.1 / 1\.2
$1\.25
$4\.25
$5\.50
[Meta](https://dev.meta.ai/docs/pricing-rate-limits)
**GLM\-5\.3**
**$1\.40**
**$4\.40**
**$5\.80**
[**Z\.AI**](https://docs.z.ai/guides/overview/pricing)
Grok 4\.6 — <200K prompt tokens
$2\.00
$6\.00
$8\.00
[xAI](https://docs.x.ai/developers/models/grok-4.6)
MiMo\-V2\.5 Pro \(\>256K\)
$2\.00
$6\.00
$8\.00
[Xiaomi](https://platform.xiaomimimo.com/docs/en-US/pricing)
Qwen3\.8\-Max
$2\.00
$6\.00
$8\.00
[QwenCloud](https://www.qwencloud.com/models/qwen3.8-max)
Gemini 3\.6 Flash — starting Jan\. 1, 2027
$1\.50
$7\.50
$9\.00
[Google](https://ai.google.dev/gemini-api/docs/pricing)
Gemini 3\.7 Flash — starting Jan\. 1, 2027
$1\.50
$7\.50
$9\.00
[Google](https://ai.google.dev/gemini-api/docs/pricing)
GPT\-5\.6 Terra
$2\.00
$12\.00
$14\.00
[OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)
Grok 4\.6 — ≥200K prompt tokens
$4\.00
$12\.00
$16\.00
[xAI](https://docs.x.ai/developers/models/grok-4.6)
GPT\-5\.4
$2\.50
$15\.00
$17\.50
[OpenAI](https://openai.com/api/pricing/)
Kimi K3
$3\.00
$15\.00
$18\.00
[Moonshot AI](https://platform.kimi.ai/docs/pricing/chat-k3)
Claude Opus 5
$5\.00
$25\.00
$30\.00
[Anthropic](https://platform.claude.com/docs/en/about-claude/pricing)
Sakana Fugu Ultra \(≤272K\)
$5\.00
$30\.00
$35\.00
[Sakana AI](https://console.sakana.ai/pricing#subscription-plan)
GPT\-5\.6 Sol — Standard mode
$5\.00
$30\.00
$35\.00
[OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)
Claude Fable 5 / Claude Mythos 5
$10\.00
$50\.00
$60\.00
[Anthropic](https://platform.claude.com/docs/en/about-claude/models/overview)
GPT\-5\.6 Sol — Fast mode
$10\.00
$60\.00
$70\.00
[OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)
Using the simple VentureBeat comparison of one million input tokens plus one million output tokens, GLM\-5\.3 comes to $5\.80, versus $8 for[Grok 4\.6](https://venturebeat.com/technology/spacexai-debuts-grok-4-6-overtaking-kimi-k3s-performance-and-matching-gpt-5-6-sol-for-worlds-third-best-on-artificial-analysis)at its lower context rate, $18 for[Kimi K3,](https://venturebeat.com/technology/kimi-k3s-full-weights-are-here-but-theyre-open-with-a-caveat-what-enterprises-should-know)$30 for[Claude Opus 5](https://venturebeat.com/orchestration/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows)and $35 for[GPT\-5\.6 Sol](https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov)\.
That is not a workload\-cost estimate — real bills depend heavily on the input/output mix, caching and token consumption — but it makes the relative API price tier easy to see\.
GLM\-5\.3 is not the cheapest capable model available\. Google’s current introductory price for Gemini 3\.7 Flash is $0\.75 per million input tokens and $3\.75 per million output tokens through Dec\. 31, 2026, while OpenAI’s GPT\-5\.6 Luna is priced at $0\.20 input and $1\.20 output\.
Still, Z\.ai’s price puts GLM\-5\.3 into a notably lower cost band than the premium frontier models it is increasingly benchmarked against\.
That comparison has become more relevant following the latest independent results\.[Artificial Analysis gives GLM\-5\.3 a score of 60 on its Intelligence Index](https://x.com/ArtificialAnlys/status/2089830890709135426/photo/1), tying Kimi K3 as the top performing open weights model in the world, and scoring seven points higher than GLM\-5\.2\.
Its analysis also estimates GLM\-5\.3 at about $0\.68 per Intelligence Index task, versus roughly $0\.44 for GLM\-5\.2, despite the identical API token prices\.
The difference underscores an important caveat in headline API pricing: Artificial Analysis found GLM\-5\.3 more verbose than its predecessor, so flat per\-token rates do not necessarily mean flat costs for a completed workload\.
For developers, though, the immediate change is straightforward: GLM\-5\.3 is now callable through Z\.ai’s API at the same $1\.40/$4\.40 per\-million\-token rate as GLM\-5\.2, giving teams another relatively low\-cost option for testing frontier\-class coding and agent workloads\.