GLM-5.3 hits the API at 1.4/4.4 per million tokens (2 minute read)

TLDR AI Models

Summary

GLM-5.3, a new frontier open-source language model from Chinese startup Z.ai, has launched its API at $1.40 per million input tokens and $1.40 per million input tokens and $4.40 per million output tokens, with plans for open-weight release. The model offers advanced coding capabilities and competitive pricing compared to other leading AI models.

The API for GPM-5.3 is now available. Z.ai plans to make the model's weights openly available, but it has yet to set a release date. API pricing remains unchanged from GLM-5.2, so developers gain substantially stronger coding and long-horizon agent performance without paying more.
Original Article
View Cached Full Text

Cached at: 08/19/26, 03:34 PM

# GLM-5.3 hits the API at $1.4/$4.4 per million tokens Source: [https://venturebeat.com/ai/glm-5-3-hits-the-api-at-1-4-4-4-per-million-tokens](https://venturebeat.com/ai/glm-5-3-hits-the-api-at-1-4-4-4-per-million-tokens) After a[stunning debut last week](https://venturebeat.com/technology/glm-5-3-is-here-with-advanced-cyber-capabilities-and-reportedly-already-found-a-serious-vulnerability-in-cursor)with cyber capabilities so advanced they reportedly found a previously undetected vulnerability in Cursor, GLM\-5\.3, the new frontier open source language model from Chinese startup z\.ai, has[now hit the application programming interface \(API\)](https://x.com/Zai_org/status/2089816129011098048?s=20)— allowing developers the ability to build atop it and plug it into their agents and applications\. Developers who previously subscribed to a GLM Coding Plan are currently limited to the OpenAI Chat Completions\-compatible protocol\. Z\.ai said it plans to make the model's weights openly available, but a precise date and licensing remain to be seen\. On the API, the price is unchanged from**GLM\-5\.2: $1\.40 per million input tokens and $4\.40 per million output tokens**\. Cached input costs $0\.26 per million tokens, while Z\.ai currently lists cached\-input storage as free for a limited time\. That means developers can move to the new generation without taking a higher posted per\-token rate from Z\.ai, even as the company claims substantially stronger coding and long\-horizon agent performance\. At those rates, GLM\-5\.3 sits well below several of the highest\-end frontier APIs\. **Model** **Input \($/1M\)** **Output \($/1M\)** **Total \($/1M\)** **Source** Muse Spark 1\.2 Contributor $0\.10 $0\.20 $0\.30 [Meta](https://dev.meta.ai/docs/pricing-rate-limits) MiMo\-V2\.5 Flash $0\.10 $0\.30 $0\.40 [Xiaomi](https://platform.xiaomimimo.com/docs/en-US/pricing) DeepSeek\-V4\-Flash — off\-peak $0\.22 $0\.66 $0\.88 [DeepSeek](https://x.com/deepseek_ai/status/2087864589895798968) GPT\-5\.6 Luna $0\.20 $1\.20 $1\.40 [OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/) MiniMax\-M3 $0\.30 $1\.20 $1\.50 [MiniMax](https://platform.minimax.io/subscribe/token-plan?tab=api-enterprise) LongCat\-2\.0 — limited\-time promo $0\.30 $1\.20 $1\.50 [LongCat](https://longcat.chat/platform/docs/APIPayAsYouGo.html) DeepSeek\-V4\-Flash — peak hours $0\.44 $1\.32 $1\.76 [DeepSeek](https://x.com/deepseek_ai/status/2087864589895798968) MiMo\-V2\.5 $0\.40 $2\.00 $2\.40 [Xiaomi](https://platform.xiaomimimo.com/docs/en-US/pricing) DeepSeek\-V4\-Pro — off\-peak $0\.66 $1\.98 $2\.64 [DeepSeek](https://x.com/deepseek_ai/status/2087864589895798968) LongCat\-2\.0 — standard $0\.75 $2\.95 $3\.70 [LongCat](https://longcat.chat/platform/docs/APIPayAsYouGo.html) MiMo\-V2\.5 Pro \(≤256K\) $1\.00 $3\.00 $4\.00 [Xiaomi](https://platform.xiaomimimo.com/docs/en-US/pricing) Gemini 3\.6 Flash — through Dec\. 31, 2026 $0\.75 $3\.75 $4\.50 [Google](https://ai.google.dev/gemini-api/docs/pricing) Gemini 3\.7 Flash — through Dec\. 31, 2026 $0\.75 $3\.75 $4\.50 [Google](https://ai.google.dev/gemini-api/docs/pricing) DeepSeek\-V4\-Pro — peak hours $1\.32 $3\.96 $5\.28 [DeepSeek](https://x.com/deepseek_ai/status/2087864589895798968) Muse Spark 1\.1 / 1\.2 $1\.25 $4\.25 $5\.50 [Meta](https://dev.meta.ai/docs/pricing-rate-limits) **GLM\-5\.3** **$1\.40** **$4\.40** **$5\.80** [**Z\.AI**](https://docs.z.ai/guides/overview/pricing) Grok 4\.6 — <200K prompt tokens $2\.00 $6\.00 $8\.00 [xAI](https://docs.x.ai/developers/models/grok-4.6) MiMo\-V2\.5 Pro \(\>256K\) $2\.00 $6\.00 $8\.00 [Xiaomi](https://platform.xiaomimimo.com/docs/en-US/pricing) Qwen3\.8\-Max $2\.00 $6\.00 $8\.00 [QwenCloud](https://www.qwencloud.com/models/qwen3.8-max) Gemini 3\.6 Flash — starting Jan\. 1, 2027 $1\.50 $7\.50 $9\.00 [Google](https://ai.google.dev/gemini-api/docs/pricing) Gemini 3\.7 Flash — starting Jan\. 1, 2027 $1\.50 $7\.50 $9\.00 [Google](https://ai.google.dev/gemini-api/docs/pricing) GPT\-5\.6 Terra $2\.00 $12\.00 $14\.00 [OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/) Grok 4\.6 — ≥200K prompt tokens $4\.00 $12\.00 $16\.00 [xAI](https://docs.x.ai/developers/models/grok-4.6) GPT\-5\.4 $2\.50 $15\.00 $17\.50 [OpenAI](https://openai.com/api/pricing/) Kimi K3 $3\.00 $15\.00 $18\.00 [Moonshot AI](https://platform.kimi.ai/docs/pricing/chat-k3) Claude Opus 5 $5\.00 $25\.00 $30\.00 [Anthropic](https://platform.claude.com/docs/en/about-claude/pricing) Sakana Fugu Ultra \(≤272K\) $5\.00 $30\.00 $35\.00 [Sakana AI](https://console.sakana.ai/pricing#subscription-plan) GPT\-5\.6 Sol — Standard mode $5\.00 $30\.00 $35\.00 [OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/) Claude Fable 5 / Claude Mythos 5 $10\.00 $50\.00 $60\.00 [Anthropic](https://platform.claude.com/docs/en/about-claude/models/overview) GPT\-5\.6 Sol — Fast mode $10\.00 $60\.00 $70\.00 [OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/) Using the simple VentureBeat comparison of one million input tokens plus one million output tokens, GLM\-5\.3 comes to $5\.80, versus $8 for[Grok 4\.6](https://venturebeat.com/technology/spacexai-debuts-grok-4-6-overtaking-kimi-k3s-performance-and-matching-gpt-5-6-sol-for-worlds-third-best-on-artificial-analysis)at its lower context rate, $18 for[Kimi K3,](https://venturebeat.com/technology/kimi-k3s-full-weights-are-here-but-theyre-open-with-a-caveat-what-enterprises-should-know)$30 for[Claude Opus 5](https://venturebeat.com/orchestration/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows)and $35 for[GPT\-5\.6 Sol](https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov)\. That is not a workload\-cost estimate — real bills depend heavily on the input/output mix, caching and token consumption — but it makes the relative API price tier easy to see\. GLM\-5\.3 is not the cheapest capable model available\. Google’s current introductory price for Gemini 3\.7 Flash is $0\.75 per million input tokens and $3\.75 per million output tokens through Dec\. 31, 2026, while OpenAI’s GPT\-5\.6 Luna is priced at $0\.20 input and $1\.20 output\. Still, Z\.ai’s price puts GLM\-5\.3 into a notably lower cost band than the premium frontier models it is increasingly benchmarked against\. That comparison has become more relevant following the latest independent results\.[Artificial Analysis gives GLM\-5\.3 a score of 60 on its Intelligence Index](https://x.com/ArtificialAnlys/status/2089830890709135426/photo/1), tying Kimi K3 as the top performing open weights model in the world, and scoring seven points higher than GLM\-5\.2\. Its analysis also estimates GLM\-5\.3 at about $0\.68 per Intelligence Index task, versus roughly $0\.44 for GLM\-5\.2, despite the identical API token prices\. The difference underscores an important caveat in headline API pricing: Artificial Analysis found GLM\-5\.3 more verbose than its predecessor, so flat per\-token rates do not necessarily mean flat costs for a completed workload\. For developers, though, the immediate change is straightforward: GLM\-5\.3 is now callable through Z\.ai’s API at the same $1\.40/$4\.40 per\-million\-token rate as GLM\-5\.2, giving teams another relatively low\-cost option for testing frontier\-class coding and agent workloads\.

Similar Articles

GLM-5.2 is probably the most powerful text-only open weights LLM

Simon Willison's Blog

Chinese AI lab Z.ai released GLM-5.2, a 753B parameter open weights LLM with a 1M token context window under MIT license, achieving top scores on the Artificial Analysis Intelligence Index and ranking second on the Code Arena WebDev leaderboard.

GLM-5.2 is a step change for open agents

Hacker News Top

Z.ai released GLM-5.2, an open-weight AI model that represents a step change for open agents, with strong benchmark performance and community hype, positioning it as the only open model competing with top closed models from OpenAI and Anthropic.

GLM-5.2 (6 minute read)

TLDR AI

Z.ai launched GLM-5.2 with a 1 million-token context window, new reasoning controls, and support for long-horizon coding tasks. It is available immediately to Coding Plan users, with API access, chatbot support, and MIT-licensed open weights coming next week.

zai-org/GLM-5.2 is here!

Reddit r/LocalLLaMA

Z.AI releases GLM-5.2, a new flagship model with a solid 1M-token context, enhanced coding capabilities with flexible thinking effort, and improved architecture via IndexShare. It is released under an MIT open-source license.

GLM-5.2 is the new leading open weights model on Artificial Analysis

Hacker News Top

Z ai's GLM-5.2 has become the new leading open weights model on the Artificial Analysis Intelligence Index, scoring 51 and outperforming competitors like MiniMax-M3 and DeepSeek V4 Pro. The model features 744B total parameters, 40B active, MIT license, and 1M context window.