Z.ai released GLM 5.3 on August 14, 2026, and the launch is unusual in one very specific way. This is not a new architecture and not a new pretrain. GLM 5.3 runs on the same 743-billion-parameter base as its predecessor, and every reported gain comes from extended post-training. By the company’s own figures it now leads CyberGym at 84.5 percent, ahead of both Claude Mythos 5 and GPT-5.6 Sol.

It also shipped without the open weights that made the GLM line matter in the first place. Z.ai says API access and weights will arrive in stages after a safety evaluation, roughly two weeks out. That decision is the real story here. Below is what actually changed, what the benchmark chart quietly leaves out, why Z.ai says it held the weights back, and how to reach GLM 5.3 in the meantime.

The Key Takeaways

  • Released: August 14, 2026, post-trained on the same 743B base as GLM 5.2. No new pretrain.
  • Biggest jump: Terminal-Bench 3.0 went from 4.6 to 28.3, and DeepSWE v1.1 from 46.2 to 66.9.
  • Security lead: CyberGym rose to 84.5 percent, edging Mythos 5 at 83.8 and GPT-5.6 Sol at 83.6, according to Z.ai’s own chart.
  • No weights yet: a break from GLM 5.2, which shipped MIT-licensed weights within days. Z.ai promises them in about two weeks.
  • Available now: GLM Coding Plan and ZCode only. No public per-token API price has been published for 5.3.

What GLM 5.3 Actually Is

De l'éditeur

Tous les modèles d'IA dans une seule app

Fello AI réunit GPT-5.6, Claude 5, Gemini 3.6, Grok 4.5 et plus dans une seule app native pour Mac et iPhone.

Téléchargez maintenant !

GLM 5.3 is Z.ai’s coding and agentic model, and the headline technical fact is what did not change. The base model is identical to the one underneath GLM 5.2 and its 1M-token context. Z.ai did not retrain it. Instead the company scaled post-training: more task environments, more environment types, and longer training runs on top of the existing 743B foundation.

That matters more than it sounds. For most of the past two years, a version bump in a frontier model line meant a bigger or freshly trained base. GLM 5.3 shows a lab pulling a large capability jump out of post-training alone, on hardware it has already paid for. New to this model family? The background on what GLM actually is covers how Zhipu’s open line got here.

Z.ai framed the launch on X in three lines, and the framing is worth reading closely, because “cyber defense” is doing a lot of work in it.

GLM 5.3 Benchmarks: the Wins and the Losses

Every number below comes from Z.ai’s own launch chart. There is no independent audit of GLM 5.3 yet, and the comparison figures for rival models are the ones Z.ai chose to publish.

Read them as a vendor’s best case, not a neutral scoreboard.

BenchmarkGLM 5.3GLM 5.2What it measures
Terminal-Bench 3.028.34.6Long-horizon terminal tasks
DeepSWE v1.166.946.2Real software issue resolution
Agents’ Last Exam CLI28.523.8Agentic command-line work
CyberGym84.5%77.2%Vulnerability discovery and validation
ExploitBench54.4%24.4%Exploitation capability
ExploitGym (2hr / 6hr)105 / 13029 / 39Tasks completed under a time budget

Where GLM 5.3 leads

The Terminal-Bench result is the one to look at.

Going from 4.6 to 28.3 is not an incremental gain, it is a model that could barely hold a long terminal session becoming one that can. DeepSWE v1.1 climbing twenty points tells a similar story about sustained, multi-step work rather than single-turn answers.

On security, CyberGym at 84.5 percent puts GLM 5.3 fractionally ahead of the closed frontier models Z.ai lists. That is a genuine milestone for an open-weight line, with the obvious caveat that the weights are not out yet.

Where it still trails

The chart is less flattering elsewhere, and most launch coverage skipped this part.

On Z.ai’s own internal Code Bench, GLM 5.3 reaches 31.4 percent at a 50,000-token budget against Claude Opus 4.8 at 29.5 percent, but Claude Fable 5 leads the same test at 39.5 percent. On ExploitBench, the jump from 24.4 to 54.4 is dramatic in relative terms and still well behind the frontier leaders, which Z.ai reports at 78.0 percent.

So the honest summary is narrower than the headlines: GLM 5.3 is at or near the frontier on vulnerability discovery and long-horizon agentic coding, and clearly behind it on raw code quality and on exploitation. For a wider view of how this line compares across the board, see how GLM stacks up against Claude, GPT and Gemini.

The Capability That Outgrew Its Training

The most interesting disclosure in this launch is not a benchmark. Z.ai added vulnerability discovery environments to post-training expecting the model to get better at spotting individual bugs. According to the company, it got something else. Capability “continued compounding as training scaled, and the model began reasoning across multiple stages of exploitation, forming coherent plans for complete exploitation chains.”

In plain terms, Z.ai aimed for a better bug finder and produced a model that plans across a whole attack chain. The company also reports real-world output from this line: 2,436 vulnerabilities identified across 269 open-source projects since GLM 5.2, with 1,097 rated critical or high severity. Those are defensive findings, and they are the strongest argument for the “cyber defense” framing.

They are also, unavoidably, evidence of the same capability pointed the other way.

This is not the first time in 2026 that a lab has described exactly this problem. It is worth reading alongside the running record of AI safety incidents in 2026, and the case where AI models attacked Hugging Face.

Why the Open Weights Are Not Out

GLM 5.2 shipped MIT-licensed weights to Hugging Face within days of launch. That was the whole proposition: frontier-adjacent performance you could download, inspect and self-host. GLM 5.3 broke that pattern on day one. As of publication there is no GLM 5.3 repository on Z.ai’s Hugging Face organisation, where GLM 5.2 is still the most recent release listed.

Z.ai’s stated reason is the capability described above. The company says API access and open weights will be released in stages following safety evaluations, and reporting from Unite.AI puts that at roughly two weeks, so the end of August 2026.

If that sounds familiar, it should. OpenAI made a structurally identical decision this month, which we covered in OpenAI’s hacking model and why you cannot use it. Two labs, two weeks apart, both concluded that a cyber-capable model needed gating before general release.

That is a pattern, not a coincidence. It has real consequences for anyone whose roadmap assumes downloadable frontier models, and the current state of play is in our guide to the best open-source AI models.

One caveat on the promise itself. A two-week timeline announced on launch day is a stated intention, not a shipped artefact. GLM 5.2’s weights are downloadable today. GLM 5.3’s are not, and the only thing standing behind the date is Z.ai’s word.

How to Use GLM 5.3 Today

Access is narrow at launch. Z.ai’s follow-up post states that GLM 5.3 is available through the GLM Coding Plan and ZCode, with broader API access staged behind the same safety review. Notably, the public subscribe page still advertises the plan as powered by GLM 5.2 and GLM-5-Turbo, so the rollout is visibly mid-flight.

Reported Coding Plan pricing runs at three tiers: Lite at $18 a month, Pro at $80 and Max at $168, with annual billing bringing those to effective monthly rates of $12.60, $56 and $117.60.

No per-token rate has been published for GLM 5.3 at all. Anyone budgeting API spend is still working from the GLM 5.2 table. Our breakdown of GLM pricing and the Coding Plan tiers covers how that quota system behaves in practice.

Using GLM 5.3 on a Mac

ZCode is Z.ai’s own coding agent, and it is optional rather than mandatory. The Coding Plan itself works with Claude Code, Cline, Roo Code and a long list of other harnesses, which is the path most Mac developers will want. There is no weight download to wait on for any of those routes, because the model runs server-side either way.

There is another option if you would rather not buy a coding subscription just to try one model. Fello AI already puts GLM alongside Claude, GPT, Gemini and Grok in a single native Mac app, which is a reasonable place to sit while the 5.3 weights clear review. The GLM desktop client for macOS page covers that setup.

Is GLM 5.3 Worth Switching To

If you are already on the GLM Coding Plan, this is a free upgrade to a meaningfully better agentic coder, and the Terminal-Bench and DeepSWE jumps are large enough to notice in daily use.

Take it.

If you were waiting on GLM 5.3 specifically to self-host a frontier-class open model, wait. There is nothing to download, and the value of this line has always been the licence as much as the scores.

Check back at the end of August. Until then GLM 5.2 remains the strongest thing you can actually run, and our notes on what is rumoured for GLM 5.5 track what comes after.

And if you are comparing coding models on merit rather than licence, GLM 5.3 has not displaced Claude at the top of the code-quality tests, by Z.ai’s own admission. The verdict for now: an excellent agentic coder, a genuine security milestone, and an open model that is not yet open.

FAQ

Is GLM 5.3 open source?

Not at launch. Z.ai says open weights and wider API access will be released in stages after safety evaluation, roughly two weeks from the August 14, 2026 release. That breaks the GLM 5.2 pattern, which shipped MIT-licensed weights within days. There is no GLM 5.3 repository on Hugging Face yet.

How much does GLM 5.3 cost?

It is included in every GLM Coding Plan tier, reported at $18, $80 and $168 a month, or $12.60, $56 and $117.60 on annual billing. Z.ai has not published a per-token API price for GLM 5.3, so the public rate card still shows GLM 5.2.

Is GLM 5.3 better than Claude for coding?

On Z.ai’s own Code Bench it edges Claude Opus 4.8 while using far fewer tokens, but Claude Fable 5 still leads that test at 39.5 percent. GLM 5.3 is stronger on long-horizon agentic tasks and vulnerability discovery, and behind on raw code quality.

Can I run GLM 5.3 on a Mac?

You can use it on a Mac, but not run it locally, because no weights have been published. The GLM Coding Plan works with Claude Code, Cline and other harnesses that run on macOS, and Z.ai’s own ZCode agent is optional. Any multi-model Mac client that supports GLM will reach it server-side.

What is CyberGym?

CyberGym is a benchmark that tests a model’s ability to discover and validate software vulnerabilities from white-box source code. Z.ai reports GLM 5.3 at 84.5 percent, up from 77.2 percent for GLM 5.2, which it says places the model ahead of Claude Mythos 5 and GPT-5.6 Sol.