GLM pricing is the headline reason developers keep switching to Zhipu AI’s models. The top-end GLM-5.3 costs just $1.40 per million input tokens and $4.40 per million output on the official API, which used to be a sixth of what GPT-5.5 charged, and the older GLM-4.7 is cheaper still. OpenAI has cut its own rates twice since then, so the gap today is closer to half. You can also use GLM free in the browser, and three of the models cost nothing at all through the API. GLM 5.3 carries the same rate card as the GLM-5.2 it replaced, so a version bump does not change your bill.
This guide breaks down every way you pay for GLM in 2026: the free tier, the per-token API rates for each model, and the GLM Coding Plan subscription that starts at $18 a month. We finish with a cost comparison against Claude, ChatGPT, Gemini and DeepSeek, updated for OpenAI’s GPT-6 ladder, where Luna at $0.10 / $0.50 undercuts every paid GLM model outright, plus a flat-rate alternative if you would rather skip per-token math entirely. Price is only half the story, so see how GLM stacks up against the top closed models on quality too. If you are new to the family, start with what GLM is first.
The Key Takeaways
- GLM is free to use at chat.z.ai, and GLM-4.7 Flash, GLM-4.5 Flash and the vision model GLM-4.6V Flash are completely free on the API too.
- GLM-5.3 API pricing is $1.40 in / $4.40 out per million tokens, less than half of GPT-6.1 Sol on output and a fifth of Claude Opus 5.5, though GPT-6 Luna is cheaper than both.
- GLM-4.7 is the budget coding pick at $0.60 in / $2.20 out per million tokens, with cached input as low as $0.11.
- The GLM Coding Plan is a flat subscription at $18 (Lite), $80 (Pro), and $168 (Max) per month, or 30% less on annual billing.
- GLM-5.3 is on the rate card at the same $1.40 / $4.40 as GLM-5.2, and since 30 July 2026 the Coding Plan is metered in credits rather than prompts and no longer serves GLM-5.2 at all.
- Inside Fello AI, GLM-5.3 is bundled with Claude, GPT and Gemini for one flat $9.99 a month, with no per-token billing.
GLM Pricing at a Glance
Here is the whole GLM pricing picture in one table. The free options are genuinely free, the API rates undercut every closed flagship except OpenAI’s cheapest tier, and the Coding Plan is a flat monthly fee for heavy use.
| What you pay for | Cost |
|---|---|
| GLM chat (chat.z.ai) | Free |
| GLM-4.7 Flash / GLM-4.5 Flash / GLM-4.6V Flash API | Free |
| GLM-5.3 API | $1.40 in / $4.40 out per million tokens |
| GLM-5.3 Flash API | $0.15 in / $0.50 out per million tokens |
| GLM-4.7 API | $0.60 in / $2.20 out per million tokens |
| GLM Coding Plan | $18 / $80 / $168 per month |
Is GLM Free?
Yes, in more ways than most rivals. You can chat with GLM-5.3 for free at chat.z.ai with no subscription, the same way you would use a free ChatGPT account. Every GLM up to GLM-5.2 is open-weight and MIT-licensed, so you can download the weights from Hugging Face and run them yourself at zero licensing cost. GLM-5.3’s weights are public too, on Hugging Face since 28 August 2026, but under a bespoke GLM 5.3 License rather than MIT, and at 753B parameters they need a multi-GPU server rather than a laptop.
The surprise is that the API has free models too. GLM-4.7 Flash, GLM-4.5 Flash and the vision model GLM-4.6V Flash are listed at $0 for input, cached input and output, so you can build real apps on them without paying per token. That makes GLM one of the cheapest serious models to prototype on, and it is a big part of why it lands on lists of the best open source AI models.
GLM API Pricing by Model
On the official Z.ai API, every model is billed per million tokens, split into input, cached input and output. Cached input applies when you reuse the same context, and it cuts the input rate to roughly a fifth. The full per-model rates are below, taken from Z.ai’s pricing docs.
| Model | Input / M | Cached input / M | Output / M |
|---|---|---|---|
| GLM-5.3 | $1.40 | $0.26 | $4.40 |
| GLM-5.3 Flash | $0.15 | $0.03 | $0.50 |
| GLM-4.7 | $0.60 | $0.11 | $2.20 |
| GLM-4.7 Flash | Free | Free | Free |
| GLM-4.5 Air | $0.20 | $0.03 | $1.10 |
GLM-5.3 (Flagship Rate)
GLM-5.3 is the model most people mean by GLM pricing today, and the GLM whose rate almost everyone is quoting. At $1.40 input and $4.40 output per million tokens, it costs less than half of GPT-6.1 Sol on output and a fifth of Claude Opus 5.5’s $20.00 output rate. Cached input drops to $0.26, so long, repeated contexts get even cheaper. Worth knowing before you pitch it as the cheap option: on Artificial Analysis’s Intelligence Index v4.3.2, GPT-6.1 Sol scores 52 to GLM-5.3’s 45 and costs $0.72 per index task against GLM’s $2.01, so the case for comparable quality at a fraction of the price reverses against OpenAI’s mid tier. Full benchmarks live in our GLM 5.3 breakdown.
GLM-5.2 (Same Rate, Older Model)
GLM-5.2 is still on the rate card at exactly the same $1.40 / $0.26 / $4.40, so the two generations cost the same and Z.ai’s migration guide simply tells developers to point the model identifier at glm-5.3. On the GLM Coding Plan you cannot choose it any more: since 30 July 2026, requests for GLM-5.2 and GLM-5.1 are routed to GLM-5.3, and GLM-4.7 to GLM-5.3 Flash. What GLM-5.2 keeps is its licence, it is the largest GLM published under MIT, and its own numbers, which sit in our GLM 5.2 breakdown. Its cheaper sibling GLM 5.3 Flash arrived on August 26, 2026 at $0.15 input and $0.50 output per million tokens, and the promotion that halved those rates ended on September 9.
GLM-5.3 Flash and FlashX
GLM-5.3 Flash is the cheap, natively multimodal sibling at $0.15 input and $0.50 output per million tokens, with GLM-5.3 FlashX sitting between it and the flagship at $0.37 / $1.25. Flash scores 42 on Artificial Analysis’s Intelligence Index v4.3.2 against the flagship’s 45, at $0.25 per index task rather than $2.01, which makes it the better value of the two for most high-volume work.
GLM-4.7 (Budget Coding Pick)
GLM-4.7 is the value champion at $0.60 input and $2.20 output per million tokens, with cached input as low as $0.11. It still scores 73.8% on SWE-bench Verified, one of the best results any open model has posted, so you get serious coding for a fraction of frontier prices. On the Coding Plan you no longer reach it directly, though: since 30 July 2026, GLM-4.7 requests are routed to GLM-5.3 Flash.
GLM-4.7 Flash (Free)
GLM-4.7 Flash is completely free on the API, for input, cached input and output alike. It is a smaller, faster model aimed at local and high-volume coding, and alongside GLM-4.5 Flash it lets you ship a working product without paying per token.
GLM-4.5 Air
GLM-4.5 Air is the lightweight legacy option at $0.20 input and $1.10 output per million tokens, with cached input at just $0.03. It is not the cheapest paid model any more, GLM-4.7-FlashX at $0.07 / $0.40 and GLM-5.3 Flash at $0.15 / $0.50 both come in under it, but it still handles everyday chat and reasoning well.
The GLM Coding Plan: Lite, Pro, and Max
If you code all day, per-token billing gets fiddly, so Zhipu sells a flat-rate GLM Coding Plan instead. It is the subscription that went viral with developers hunting for a cheaper Claude alternative. Since 30 July 2026 it serves only GLM-5.3 and GLM-5.3 Flash, routing GLM-5.2 and GLM-5.1 requests to GLM-5.3 and GLM-4.7 to GLM-5.3 Flash, and it meters usage in credits rather than prompts. The three tiers, from Z.ai’s subscribe page, are below.
| Tier | Monthly | Annual billing (-30%) | Usage allowance |
|---|---|---|---|
| Lite | $18 | $12.60 | 2,000 credits / 5 hrs, 10,000 / week |
| Pro | $80 | $56 | 12,000 credits / 5 hrs, 60,000 / week |
| Max | $168 | $117.60 | 28,000 credits / 5 hrs, 140,000 / week |
All three tiers give you the same model lineup and run inside coding tools like Claude Code, Cline, OpenCode and 20-plus other clients. Credit cost varies by model: GLM-5.3 bills at multipliers of 6.9 on input, 1.7 on cached input and 24 on output, GLM-5.3 Flash at 2.3, 0.56 and 8, and off-peak usage costs half the standard credit rate, with peak defined as Monday to Friday, 14:00 to 18:00 Singapore time. From 25 September to 7 October 2026 Z.ai is charging the off-peak rate around the clock. Yearly billing knocks the prices down to $151.20, $672 and $1,411.20 a year.
How to Cut Your GLM Bill
GLM is already cheap, but three habits push the cost lower. First, lean on cached input. Reusing the same system prompt or document context bills at the cached rate, which is roughly a fifth of the standard input price. Second, match the model to the job, since GLM-4.7 or the free Flash models handle most coding and chat without touching the flagship.
Third, watch the clock. The Coding Plan charges half the credit rate off-peak, and the standard API still costs far less than the flagship US tiers at every hour. That gap no longer holds against every US model, though: GPT-6 Luna at $0.10 / $0.50 is cheaper per token than any paid GLM model, GLM-4.5 Air included. If your workload is bursty, a per-token plan can beat a subscription; if you code steadily all day, the flat Coding Plan usually wins.
GLM vs Claude, ChatGPT, Gemini, and DeepSeek on Cost
GLM’s whole pitch is frontier-class output without frontier prices. The table below lines up the GLM flagships against the models you probably already pay for, using standard API rates per million tokens. GLM no longer owns the price floor: DeepSeek and GPT-6 Luna both come in under it.
| Model | Input / M | Output / M |
|---|---|---|
| GLM-5.3 (Zhipu) | $1.40 | $4.40 |
| GLM-4.7 (Zhipu) | $0.60 | $2.20 |
| Claude Opus 5.5 | $4.00 | $20.00 |
| Claude Sonnet 5.5 | $2.00* | $10.00* |
| GPT-6.1 Sol | $2.00 | $10.00 |
| GPT-6 Astra | $10.00 | $50.00 |
| GPT-6 Luna | $0.10 | $0.50 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
| DeepSeek V4 Pro | $0.66** | $1.98** |
| GLM-5.3 Flash (Zhipu) | $0.15 | $0.50 |
*Claude Sonnet 5.5, released on 28 September 2026, keeps the $2.00 / $10.00 price of Sonnet 5. **DeepSeek bills peak and off-peak: V4 Pro is $0.66 / $1.98 off-peak and $1.32 / $3.96 at peak, which runs 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.
The picture changed twice. OpenAI’s GPT-6 ladder replaced the GPT-5.6 rates this comparison used to rest on, and the new mid tier lands right on top of GLM. GLM-5.3 still costs less than half of GPT-6.1 Sol on output and just over a third of Gemini 3.1 Pro’s, but GPT-6 Luna at $0.10 / $0.50 undercuts GLM-5.3 by 14x on input and nearly 9x on output, and beats GLM-4.7 on both. Worse for the pitch, Sol is not only close on price but ahead on quality, 52 against 45 on Artificial Analysis’s Intelligence Index v4.3.2, at $0.72 per index task against GLM’s $2.01. GLM’s remaining edge is not the floor price and not comparable quality for less; it is open weights you can self-host, a flat Coding Plan, and free Flash models that Luna has no answer to. At the top of OpenAI’s ladder, GPT-6 Astra is on the rate card at $10.00 / $50.00. For a fuller picture across every major model, see our AI pricing comparison, and for the other big open-weight value pick check Qwen pricing.
A Simpler Alternative to Per-Token Billing
Per-token math is powerful but exhausting, and most people do not want to track input versus output rates or manage a separate z.ai account. If that is you, the simpler route is a flat subscription that includes GLM alongside the other frontier models.
Inside Fello, GLM-5.3 sits next to Claude, GPT and Gemini for a single $9.99 a month. You can run the same prompt through several models at once and keep whichever answer is best, with no per-token billing, no quota juggling and no China-hosted account to set up. It is the easiest way to try GLM pricing-free before you commit to the API. On a Mac it installs as a native GLM desktop client, so the flat price covers a real app rather than a browser tab.
Conclusion
GLM pricing is about as friendly as it gets for a frontier-class model. The chat is free, three API models cost nothing, GLM-5.3 runs at less than half of GPT-6.1 Sol on output, and the GLM Coding Plan caps heavy use at a flat monthly fee. It is no longer the cheapest option per token, though, now that DeepSeek and GPT-6 Luna both sit below it.
For most people the smart move is to start free at chat.z.ai, reach for GLM-4.7 or GLM-5.3 Flash when you need cheap coding via API, and step up to GLM-5.3 for the flagship. And if you would rather skip per-token billing entirely, run GLM alongside Claude and GPT in Fello for one flat price.
FAQ
How much does GLM cost?
GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output on the official Z.ai API, less than half of GPT-6.1 Sol on output. GLM-4.7 is cheaper at $0.60 / $2.20 and GLM-5.3 Flash at $0.15 / $0.50, and the chat plus three Flash models are free. GLM-5.2 is still listed at exactly the same rate as GLM-5.3.
Is GLM free to use?
Yes. You can chat with GLM free at chat.z.ai, the model weights up to GLM-5.2 are open source under the MIT license, and GLM-4.7 Flash, GLM-4.5 Flash and GLM-4.6V Flash are free on the API for input, cached input and output. GLM-5.3’s weights are public too, on Hugging Face since 28 August 2026, but under a bespoke GLM 5.3 License rather than MIT.
How much is the GLM Coding Plan?
The GLM Coding Plan has three tiers: Lite at $18, Pro at $80 and Max at $168 per month, or 30% less on annual billing. Since 30 July 2026 every tier serves GLM-5.3 and GLM-5.3 Flash only, metered in credits rather than prompts, for use in tools like Claude Code and Cline.
Is GLM cheaper than ChatGPT and Claude?
Cheaper than the flagships, yes. GLM-5.3 runs at less than half of GPT-6.1 Sol on output and a fifth of Claude Opus 5.5 ($4.00 / $20.00). It is not the cheapest option overall, though, and it is not the stronger model either: DeepSeek and GPT-6 Luna ($0.10 / $0.50) both undercut it, and GPT-6.1 Sol outscores it 52 to 45 on Artificial Analysis’s Intelligence Index v4.3.2.
What is the cheapest GLM model?
GLM-4.7 Flash, GLM-4.5 Flash and GLM-4.6V Flash are free on the API. Among paid text models the cheapest is GLM-4.7-FlashX at $0.07 input and $0.40 output per million tokens, then GLM-5.3 Flash at $0.15 / $0.50 and GLM-4.5 Air at $0.20 / $1.10.