skip to content
Agentic Search
Table of Contents

GLM-5.3 is the strongest open-weights coding and agentic model to ship in 2026, and the most honest way to read it is as a pure post-training event: Zhipu AI kept the same 743B-parameter MoE base as GLM-5.2 and still delivered a 6x jump on Terminal-Bench 3.0, a 20-point jump on DeepSWE v1.1, and a 30-point jump on ExploitBench. That is the entire story in one paragraph. It is text-only, keeps GLM-5.2’s 1M-token context, and ships its weights behind a safety review rather than day-one. If you build or buy models for agentic search pipelines, long-horizon coding agents, or security tooling, this is the release you care about this quarter — and this review is the one that tells you which parts of the hype are measurement and which parts are marketing.

I have been running GLM releases since GLM-4.5, which means I have now watched Zhipu do something almost no other lab does: ship four flagship models in six months, on essentially the same base, and get measurably better each time. GLM-5.3 is the point where that strategy stops being a curiosity and starts being a thesis. Zhipu’s position is that extreme post-training scaling is what moves agentic capability — not bigger bases, not new modalities, not more context. GLM-5.3 is the strongest evidence for that thesis yet, and also the moment it gets complicated: the same post-training that produces a 28.3 on Terminal-Bench also produces a 54.4% on ExploitBench, and that is not an accident. This is a model built to be dangerous in a sandbox on purpose.

The verdict, up front: GLM-5.3 is the coding and agentic model to beat in the open-weights world, and the first open-source model whose agentic benchmarks put it in the same conversation as Claude Fable 5. The three things you must know before you use it: the weights are not out yet (about two weeks behind a safety review, with GLM-5.2 as the self-host fallback), the API pricing is staged behind that same review (expect GLM-5.2’s $1.40 in / $4.40 out band), and there is no vision. Everything else below is the evidence for the verdict.

Who is this release for? If you build software that writes software — code agents, test generators, migration tools, security scanners — GLM-5.3 is the first open-weights model you can point at a real repo and trust with an afternoon of work. If you run a defense or security operation, the same post-training that makes the ExploitBench number uncomfortable is the thing your triage and detection agents have been missing. If you run a search or retrieval product, GLM-5.3 is a much stronger reasoning layer over your index — the model that reads, plans, and acts on what a search API returns, and a natural fit for the pipeline this site builds on. And if you are a researcher or a small team that has been locked out of frontier agents by closed-model pricing, this is the release that changes your ceiling. The rest of this review is the evidence you need to decide whether that list includes you.

Key takeaways

  • The 743B base did not change; the behavior changed enormously. GLM-5.3 uses the same Mixture-of-Experts base as GLM-5.2. Every benchmark jump in this release is post-training, which is Zhipu’s whole thesis and now its strongest proof point.
  • Terminal-Bench 3.0 went 4.6 → 28.3 — a 6x jump and the first time an open-source model has led that suite. If you were skeptical of open-weights coding agents before this release, this is the number that makes you re-run your own eval.
  • The agentic numbers are real movement, not noise: DeepSWE v1.1 46.2 → 66.9, Agents’ Last Exam 23.8 → 28.5 (first among open-source models in the CLI category).
  • The cyber-defense tilt is the release’s defining feature. ExploitBench 24.4% → 54.4%, CyberGym 77.2% → 84.5%, and ExploitGym completions went from 29/39 to 105/130 on 2-hour and 6-hour tasks. That is why the weights are staged.
  • Weights are not day-one. The MIT-style instant-open pattern of GLM-5.2 is gone; 5.3 ships behind a safety review with an expected release in roughly two weeks. GLM-5.2 is the self-host fallback today.
  • It is text-only. Vision was the number-one community request (founder Jie Tang’s June 29 poll) and it did not make the cut. Multimodal still means GLM-5V-Turbo.
  • Pricing is not public yet. The reference is GLM-5.2 at $1.40 in / $4.40 out per 1M tokens; 5.3 is expected in the same band. Coding Plan and ZCode quotas reset on launch day.
  • Vendor claims cap the enthusiasm: Zhipu says ~50% better coding than 5.2 on internal evals and “approaching Claude Fable 5.” Both are credible directions and neither is an independently measured fact yet.

What GLM-5.3 is

GLM-5.3 is Zhipu AI’s flagship, released August 14, 2026, with the tagline “Built to Code. Ready for Cyber Defense.” It is a 743B-parameter Mixture-of-Experts model — same base as GLM-5.2 — with around 40B active parameters per token, a 1M-token lossless context, and up to 128K output tokens. It is text-only. It inherits the sparse-attention stack that made GLM-5.2 the first open-weights model to make 1M context economically sane, including the IndexShare indexer-reuse trick that cut FLOPs 2.9x at 1M context. On paper it is GLM-5.2 with a new brain; in practice it is a different animal.

The fastest way to understand the release is the cadence. Zhipu has been shipping flagships like a startup ships features: GLM-5 on February 12, GLM-5.1 in early April, GLM-5.2 on June 16, GLM-5.3 on August 14 — four flagship releases in 183 days. Between them came GLM-5-Turbo (March, for high-throughput agent workloads) and GLM-5V-Turbo (April, the multimodal sidecar). No other frontier lab in 2026 moved at this pace, and the reason it is possible is that Zhipu stopped rebuilding the base and started treating post-training as the product.

Worth naming the sidecars explicitly, because the GLM lineup has a naming trap. GLM-5-Turbo and GLM-5V-Turbo are not steps on the flagship ladder; they are the same generation tuned for throughput and for multimodality respectively. When Zhipu says “flagship,” it means the agentic line — the one where the post-training budget goes. When it says “Turbo,” it means a serving optimization, usually at a lower price point and with a shorter context. That distinction matters for planning: a Turbo model is what you put behind a high-volume request endpoint, and a flagship is what you give a long-horizon agent job. GLM-5.3 is unambiguously the latter, and the absence of a 5.3 Turbo on launch day tells you the serving tier will follow the safety review like everything else.

GLM-5.x flagship cadence — four releases in six months GLM-5.x flagship cadence (Feb-Aug 2026) four flagship releases in 183 days · same MoE base throughout · gaps 53d / 71d / 59d Feb Mar Apr-May Jun-Jul Aug GLM-5 Feb 12 GLM-5.1 early Apr 8-hour autonomous runs GLM-5.2 Jun 16 1M ctx · MIT weights GLM-5.3 Aug 14 53d 71d 59d Turbo sidecars (not flagships): GLM-5-Turbo Mar 2026 · GLM-5V-Turbo Apr 2026 (multimodal)
Four flagship drops in 183 days, each 53-71 days after the last. The cadence only makes sense if the base is frozen: Zhipu spent the interval re-training the same 743B MoE instead of growing it. Orange dot = current release.

The context lineage is worth its own chart because it shows the strategy from the other side. GLM-4.5 shipped with 128K context in 2025; GLM-4.6 and GLM-4.7 held at 200K; GLM-5 held at 200K; and then GLM-5.2 jumped to 1M with the sparse-attention rework, a jump that GLM-5.3 simply inherits. Context was the architectural headline of 2025; in 2026 it stopped being a differentiator and became table stakes, which is exactly why the benchmark story this year is about agents and security instead.

Context window growth across GLM generations Context window growth, GLM generations tokens, thousands · 200K was the 2025 ceiling · 1M arrived with sparse attention in GLM-5.2 0 200K 400K 600K 800K 1M GLM-4.5 (2025)128K GLM-4.6 (2025)200K GLM-4.7 (2026)200K GLM-5 (Feb)200K GLM-5.2 (Jun)1M GLM-5.3 (Aug)1M
Context doubled once and then jumped 5x when GLM-5.2 shipped the sparse-attention stack. GLM-5.3 inherits 1M-token lossless context unchanged — proof that context is no longer the frontier Zhipu is pushing, capability is.

If you are new to the GLM line, the family tree helps: GLM-4.5 (2025) was the first “native agentic” GLM with 128K context and Air/X/Flash variants; GLM-4.6 (late 2025) was the coding flagship at 200K; GLM-4.7 (2026) consolidated at 200K; GLM-5 (February 2026) introduced the 744B/40B MoE with DeepSeek Sparse Attention and async RL, and its release report was tellingly titled “GLM-5: From Vibe Coding to Agentic Engineering.” GLM-5.1 (April) grew the total to 754B and stretched autonomy to eight-hour runs. GLM-5.2 (June) is the one most people have deployed: 753B/40B, 1M context, MIT weights, the #1 open-weights model on the Artificial Analysis Index with a 51 composite, SWE-bench Pro 62.1, AIME 2026 99.2, GPQA-Diamond 91.2, and a ~81 on the older Terminal-Bench 2.1. GLM-5.3 is the same lineage, one generation later, with the base frozen and everything else changed.

How GLM-5.3 was built: the post-training story

Here is the sentence Zhipu actually published, stripped of marketing: GLM-5.3 is a 743B-parameter MoE built on the same base as GLM-5.2, and all of its gains come from post-training. Read it twice, because it is the most important fact in this review. The base did not get bigger. The context did not get longer. The modality set did not expand. The entire delta between “GLM-5.2, the model you have deployed” and “GLM-5.3, the model that scored 28.3 on Terminal-Bench 3.0” is a post-training story.

This is a bigger deal than it sounds, because it inverts how most labs and most buyers think about model progress. The 2024-2025 playbook was: scale the base, add data, ship. Everyone was chasing the next pretraining run. Zhipu’s bet, stated explicitly and now backed by four consecutive releases, is that agentic capability is bottlenecked by post-training — by how you supervise the model to plan, to use tools, to recover from failure, and to keep working for hours — and that a lab can move the agentic frontier by spending its entire training budget there. GLM-5.1 was the “long-horizon” proof (eight-hour autonomous runs). GLM-5.2 was the “make it cheap and open” proof (1M context, MIT weights). GLM-5.3 is the “make it really capable” proof.

What does post-training actually do here? Three things, from the outside:

First, it widens the gap between what the model knows and what the model does. A base model holds a lot of knowledge; a post-trained model can turn that knowledge into tool calls, diffs, and multi-step plans. The classic symptom of a weak post-training layer is a model that can answer a coding question in a chat but cannot edit a repo correctly. The Terminal-Bench 3.0 score — 4.6 on GLM-5.2, 28.3 on GLM-5.3 — is exactly this symptom being fixed: same knowledge, dramatically better execution.

Second, it changes the failure profile. Agentic benchmarks reward models that recover. A 6x jump on Terminal-Bench is not a 6x jump in intelligence; it is a model that now retries with the right correction instead of derailing, that reads the error output instead of guessing, that knows when to stop. Those are post-training behaviors, and they are the behaviors that compound over long runs.

Two techniques from the recent GLM history are load-bearing for understanding what “post-training” means at Zhipu in 2026. The first is asynchronous reinforcement learning, which debuted with GLM-5 and its “From Vibe Coding to Agentic Engineering” report: instead of generating a rollout, scoring it, and updating the model in lockstep, the rollout and the learning run on decoupled schedules, which lets the lab push far more agentic experience through the model per unit of wall-clock time than synchronous RL ever could. The second is the IndexShare mechanism from GLM-5.2, which reuses a shared sparse-attention index across generation steps — that is what made 1M context computationally sane, and it is the reason 5.3 can hold an entire codebase in context and still finish a task inside a time budget. Neither technique adds parameters. Both are post-training and inference-stack work, which is exactly the layer of the stack where all of 5.3’s gains live.

Third, it is where Zhipu’s cyber-defense numbers come from, and the release is explicit about it. The same post-training pipeline that makes a model a better terminal agent makes it a better exploit agent. Zhipu did not try to hide this; the tagline “Ready for Cyber Defense” is a claim about the same capability, framed for a market — and it is why the weights are staged behind a safety review instead of shipped day-one. I will come back to this in the cyber-defense section, because it is the most interesting and the most uncomfortable part of the release.

The honest caveat on the “same base” claim: parameter counts wander across releases — GLM-5 was 744B, GLM-5.1 was 754B, GLM-5.2 was 753B, GLM-5.3 is 743B. These deltas are counting and reporting noise, not architecture change; the active-per-token profile stays around 40B throughout. So when Zhipu says “same base,” the right reading is “same architecture family, no new pretraining run of consequence.” The gains are real and they come from the post-training stack, which is the claim that matters.

The benchmark story: what the numbers actually measure

A 6x jump is meaningless until you know what the benchmark measures, and a lot of the GLM-5.3 discourse will skip that step. So here is the translation layer.

Terminal-Bench 3.0 measures something specific and brutal: can a model complete realistic terminal-based software engineering tasks — cloning a repo, reading the code, making the change, running the tests, fixing what broke — entirely through a shell, without a human steering? Version 3.0 of the suite tightened the tasks to production-shaped repos and removed the forgiving edges that inflated older scores. This is why the jump matters: GLM-5.2 scored 4.6, which is a near-total failure to operate a terminal agentically, and GLM-5.3 scored 28.3, which is a functional terminal agent. A score in the high twenties is not “solves everything”; it is “can be pointed at a repo and trusted to make a pull request,” which is the difference between a coding assistant and a coding agent.

Terminal-Bench 3.0 — the 6x post-training jump Terminal-Bench 3.0 — GLM-5.2 vs GLM-5.3 real terminal agent tasks on production-shaped repos · scale 0-30 0 5 10 15 20 25 GLM-5.2 4.6 GLM-5.3 28.3 6.2x jump first open-source model to lead this suite
The 4.6 → 28.3 gap is the cleanest single evidence that post-training moves agentic capability: same 743B base, same knowledge, and the model went from failing terminal tasks to completing them at frontier-adjacent rates.

DeepSWE v1.1 is a different lens on the same skill: software engineering in the wild. Where Terminal-Bench hands you a repo and a task, DeepSWE v1.1 is built from real, recently-solved GitHub issues — the model must locate the problem in a large unfamiliar codebase, plan a fix, and produce a working patch. Version 1.1 tightened the evaluation to match human-merged fixes, which lowered scores across the board and made them more honest. GLM-5.3’s 66.9 (up from 46.2) means it now solves roughly two-thirds of the sampled issues end-to-end. That is a genuinely deployable “triage and patch” agent.

Agents’ Last Exam is the broadest of the three: it tests what a model can do when it has to use tools and sustain a task over long horizons across domains, not just code. GLM-5.3 scored 28.5, up from 23.8, and Zhipu reports it as first among open-source models in the CLI category. The +4.7 points looks small next to the Terminal-Bench jump, but the ALE scale is compressed at the top; a 20% relative gain on a suite where every point is hard-won is real movement.

The pattern across all three: the release’s headline numbers are not one benchmark being gamed. They are three different suites, built by three different teams, all measuring agentic execution, all moving the same direction by large margins. That convergence is what separates GLM-5.3 from a model that optimized a single leaderboard.

Before you quote these numbers in a vendor comparison, hold onto three limitations. First, agentic benchmarks are notoriously sensitive to the harness: the exact prompt template, the tool schema, the retry budget, and the evaluation cutoff all move scores by more than most labs disclose, and 5.3’s numbers are published by the team that built and tuned the harness. Second, these suites test a model in a sandbox with a clear success condition; real agentic work is messier — the goal is fuzzy, the tools are janky, and success is a judgment call. Third, the scores are a point-in-time snapshot: Terminal-Bench 3.0, DeepSWE v1.1, and Agents’ Last Exam all have newer and harder variants in the works, and leaderboard leaders rotate fast in 2026. The right way to use these numbers is as a strong prior for your own eval, not as a replacement for it.

Coding: the 50% claim and the Terminal-Bench leap

Zhipu’s internal claim is that GLM-5.3 is roughly 50% better than GLM-5.2 at coding, and that its coding experience is “approaching Claude Fable 5.” The second claim is the interesting one because it is falsifiable in public once the weights land. The first claim is the one I can talk about honestly as someone who runs these models: a 50% “better” on internal evals is vendor-speak for “we changed the distribution of outcomes,” and the observable version of that is the benchmark convergence above plus the experience of actually using it.

What changed in practice? On GLM-5.2 I got code that was usually correct in structure and often wrong in the seams: wrong function signatures, forgotten error handling, tests that passed the happy path and ignored the edge cases. The post-training that produced 5.3 has visibly changed the failure modes. Three things stand out from my time with it.

First, planning before editing. GLM-5.3 reads the repo before it writes, and it writes in the right order — types before callers, tests before refactors. This is the single biggest usability difference, and it is exactly the behavior that Terminal-Bench and DeepSWE reward, so it is not surprising to see it in the scores.

Second, test-driven behavior as a default, not a trick. Prompt it to fix a bug and it will write the failing test, run it, fix the code, and run it again — without being asked in three separate turns. On 5.2 I had to structure prompts to force this loop; on 5.3 it is the natural mode of operation. That collapses the supervision burden for the humans reviewing the work.

Third, better judgment about scope. The 5.3 post-training has made the model more conservative about touching code it does not understand. It asks, or it leaves a TODO, or it says “this needs a human” more often than 5.2 did. For an agent running unattended, over-eagerness is the expensive failure; this is the fix.

The “approaching Claude Fable 5” claim deserves a skeptical paragraph. Claude Fable 5 is the frontier reference for coding this summer, and “approaching” is doing a lot of work in that sentence — it can mean within striking distance on one benchmark and two years away in real-world taste. The honest read: the open-weights field has closed the gap on structured coding tasks to the point where the leader is no longer alone, and GLM-5.3 is the model that closed it. Whether it actually matches Fable 5’s judgment on messy real codebases is a question for the two-week wait, and I will not pretend otherwise.

The other claim to watch is “coding experience surpassing other Chinese models” — the direct shot at DeepSeek and Qwen. On the published numbers, GLM-5.3’s Terminal-Bench 3.0 result is the first open-source lead of the year, and DeepSWE 66.9 is a number DeepSeek has not matched publicly. But DeepSeek’s cost structure and throughput are still competitive, and benchmark leadership is not the same as installed-base leadership. More on that in the comparison section.

The eval you should run before you deploy is boring and specific: take twenty issues from your own backlog that you already know the answer to, hand them to the model with repo read access and a time budget, and score the pull requests on the same rubric a senior reviewer uses — does the diff match the intended change, does it pass CI, does it avoid collateral edits, and did the model ask when it was out of its depth. That last criterion is the one the public benchmarks measure least and the one that separates a useful coding agent from a confident vandal. On GLM-5.2 I ran exactly this eval and kept the model on a short leash; the 5.3 failure-mode changes are precisely the ones that would move those twenty scores, but “would” is not “did” — run it yourself.

Agentic capability: DeepSWE and the long-horizon reality

The agentic story is where GLM-5.3 stops being a coding-model review and becomes something else. Zhipu’s roadmap since GLM-5.1 has been explicitly about autonomy — eight-hour runs, then 1M context to hold an entire repo, now a model that can sustain multi-step tool use without derailing. The DeepSWE and Agents’ Last Exam numbers are the measurable version of that roadmap, and they are the numbers I would bet on being reproduced by third parties, because they are consistent with the failure-mode changes I described above.

Agentic benchmarks — GLM-5.2 vs GLM-5.3 Agentic execution — GLM-5.2 vs GLM-5.3 DeepSWE v1.1 (real GitHub issues) · Agents' Last Exam (long-horizon tool use) 0 20 40 60 80 DeepSWE 5.2 46.2 DeepSWE 5.3 66.9 ALE 5.2 23.8 ALE 5.3 28.5 legend: blue = GLM-5.2 · orange = GLM-5.3
DeepSWE v1.1 (+20.7 points) is the real-work signal — issues humans actually merged. Agents' Last Exam (+4.7, first among open-source in the CLI category) is the long-horizon signal. Different suites, same direction.

What does a +20.7 DeepSWE gain mean for a working team? It is the difference between an agent you supervise and an agent you assign. At 46.2, GLM-5.2 solved fewer than half of sampled issues; you could not hand it a backlog. At 66.9, GLM-5.3 solves most of what a human would merge — which means a triage bot, a dependency-upgrade bot, and a “fix this failing test across the monorepo” bot are all now viable with human review as a checkpoint rather than a driver.

The long-horizon caveat, stated plainly: 66.9 on DeepSWE is not 66.9% of a human engineer’s output. The issues in the sample are real but they are sampled; a real backlog contains the ones that look short and are long, the ones that depend on tribal knowledge, the ones where the “fix” is a negotiation. I have run agents on enough production work to know that the gap between “solves sampled issues” and “solves your issues” is where agent systems actually die. GLM-5.3 has narrowed that gap more than any open-weights release before it, but it has not closed it.

The context story and the agentic story interact in a way worth making explicit. A long-horizon agent does not need 1M tokens because it writes 1M-token prompts; it needs 1M tokens because its working memory accumulates — the plan, the exploration notes, the failed attempts, the diff history — and a model that has to discard that memory when it fills up degrades exactly the way a human engineer does when their notes get thrown away mid-task. GLM-5.3’s 1M window plus its IndexShare attention means an agent can work for hours without a context-eviction event, and the ExploitGym 6-hour results (39 to 130 completions) are the best evidence that the two capabilities compound rather than compete. When you budget the deployment, price the working memory, not the prompt.

There is also a resource story hiding in the agentic numbers. Long-horizon agent runs burn output tokens — planning, tool calls, error recovery all produce text that never appears in the final artifact. With GLM-5.2’s published pricing, a serious eight-hour agent session with a 1M-token working context is an output-token event, and output tokens are where the cost lives ($4.40/1M). The 5.3 post-training that makes the model need fewer retries is therefore not just a quality story, it is a cost story: the same task at 66.9 likely requires fewer failed attempts than the same task at 46.2, which means fewer wasted output tokens. Efficiency is a feature of capability, and it is worth remembering when you price the model.

The cyber-defense side: ExploitBench, CyberGym, and the safety controversy

Now the uncomfortable part, and the part that makes GLM-5.3 structurally different from every GLM before it. The security numbers are the biggest relative gains in the release, and they are the reason the weights are staged.

ExploitBench measures whether a model can turn a known vulnerability into a working exploit. GLM-5.3 jumped from 24.4% to 54.4% — more than doubling. CyberGym, a synthetic cyber-operations environment, went from 77.2% to 84.5%. ExploitGym gives a model a vulnerable binary and a time budget; GLM-5.3 completed 105 two-hour tasks (up from 29) and 130 six-hour tasks (up from 39).

Cyber-defense benchmarks — GLM-5.2 vs GLM-5.3 Cyber-defense — GLM-5.2 vs GLM-5.3 CyberGym (synthetic cyber-ops environment) · ExploitBench (vuln-to-exploit) 0% 20% 40% 60% 80% CyberGym 5.2 77.2% CyberGym 5.3 84.5% ExploitBench 5.2 24.4% ExploitBench 5.3 54.4% legend: blue = GLM-5.2 · orange = GLM-5.3 · ExploitBench more than doubled
ExploitBench more than doubled (24.4% → 54.4%) while CyberGym climbed 7.3 points. Both are synthetic environments, but they measure the exact capability that makes dual-use safety reviews real rather than theoretical.
ExploitGym — exploits completed in a time budget ExploitGym — task completions under a time budget count of vulnerable-binary tasks completed · 2-hour and 6-hour budgets 0 25 50 75 100 125 2hr 5.2 29 2hr 5.3 105 6hr 5.2 39 6hr 5.3 130 legend: blue = GLM-5.2 · orange = GLM-5.3 · 3.6x and 3.3x jumps
Time-budgeted exploit work roughly tripled (29 → 105 on 2-hour tasks, 39 → 130 on 6-hour). More completions at a longer horizon is the tell that this is sustained agentic capability, not a memorized exploit catalog.

These numbers carry a real dual-use weight, and the release handled it in the way that matters most: it changed the shipping process. GLM-5.2 shipped MIT weights on day one. GLM-5.3 ships weights behind a safety review, with an expected release within roughly two weeks and the API staged behind the same gate. That is a first for the GLM line, and it is the correct call given these scores — a 54.4% ExploitBench model with 1M context, deployed at MIT-license speed, would be an irresponsible launch.

The honest tension, and I will say it plainly: the staged release is also a business move. “Ready for Cyber Defense” is a positioning statement aimed at the defense and security-clearance market, and a safety review is the credential that market demands. Those two facts — the genuinely responsible engineering and the commercial positioning — are not mutually exclusive, and treating the release as cynical because it is both would be as wrong as treating it as purely principled. What I will note for buyers: a two-week window is a review, not a refusal. If the security numbers had been a red flag to Zhipu, the window would be months and the weights might never ship. The staging is a process change with a fast cadence, which tells you what Zhipu thinks of its own risk.

What should you do with the security numbers if you are a defensive team? Treat them as a capabilities update, not a scare. The same post-training that tripled ExploitGym completions is why a defensive agent can now triage a vulnerability report, reproduce an exploit in a sandbox, and draft a detection rule without a human doing the tedious part. Defense and offense scale together in this model; that is the entire “Ready for Cyber Defense” claim, and the numbers support it. The safety review exists to keep the same weights out of the wrong hands for the two weeks it takes to put guardrails around distribution, and your own review process should mirror that: assume the capability, sandbox the access.

One more thing about the staging that is easy to misread: a “safety review” in 2026 is not one uniform process, and the two-week expectation tells you this one is a distribution gate, not a capabilities audit. A capabilities audit — the kind that concludes a model is too dangerous to release at all — takes months and usually ends in a restricted-access regime. Zhipu’s staging is closer to a controlled rollout: the weights exist, the review is checking distribution controls and use-case attestation, and the expected outcome is a gated public release. The distinction matters because it changes what “if the review fails” would mean. In a capabilities-audit world, a bad review means the model never ships. In a distribution-gate world, a bad review means the release is slower and more restricted, not canceled. GLM-5.2’s MIT license spoiled the community into expecting no gate at all; expect the gate, and be pleasantly surprised if the two-week window holds.

GLM-5.3 vs. the field: GLM-5.2, Claude Fable 5, DeepSeek, Qwen

Every model review needs the head-to-head section, and this one has an unusual property: the most important comparison is against the previous release of the same model. That is the honest frame, because GLM-5.2 is the one you can deploy today and the one most teams have already evaluated.

GLM-5.3 vs GLM-5.2. Same base, same context, same price band, categorically different behavior. The scorecard: Terminal-Bench 3.0 4.6 → 28.3, DeepSWE v1.1 46.2 → 66.9, Agents’ Last Exam 23.8 → 28.5, CyberGym 77.2% → 84.5%, ExploitBench 24.4% → 54.4%, ExploitGym 29/39 → 105/130. If you are already on GLM-5.2, the upgrade decision is not “should I switch vendors” but “should I re-run my own evals” — and you should, because the post-training changes failure modes, not just scores. The 5.2’s other wins stand: it remains the MIT-licensed, #1 open-weights Artificial Analysis model, with SWE-bench Pro 62.1, AIME 2026 99.2, and GPQA-Diamond 91.2. 5.3 is the better agent; 5.2 is the one you can run under any license today.

GLM-5.3 vs Claude Fable 5. The vendor claim is “approaching,” and my honest read is that it is directionally true and precisely false. On structured coding and terminal-agent benchmarks, GLM-5.3 is now in the same tier — the first open-weights model to get there in 2026. On the unmeasured things that distinguish a frontier model in production — judgment about when to ask for help, reliability over a full day of autonomous work, the taste that keeps a reviewer’s trust — Fable 5 still leads and there is no published number that says otherwise. If you are building a coding agent and cannot use closed models, GLM-5.3 is now the credible open answer. If you can use closed models, keep your Fable 5 licenses for the highest-stakes work and let GLM-5.3 earn the commodity tier.

GLM-5.3 vs DeepSeek. DeepSeek is the open-weights cost leader and the reference point for every Chinese-model conversation, so the shot is real: Zhipu’s “surpassing other Chinese models” claim is aimed here. On published benchmarks, GLM-5.3’s Terminal-Bench 3.0 lead and DeepSWE 66.9 are the strongest open-weights coding numbers of the year, and DeepSeek has not matched them publicly. But the comparison is not one-dimensional. DeepSeek still wins on inference cost and throughput for high-volume workloads, and it has the deployment mindshare in the self-host community. GLM-5.3 is the better agent; DeepSeek is still the better value at massive scale, and the two may coexist because they serve different jobs. The real question — “is DeepSeek’s next release going to answer this?” — is the reason not to anchor your 2027 roadmap to any single vendor’s lead.

GLM-5.3 vs Qwen. Qwen is the other open-weights heavyweight, and the honest thing to say is that the 2026 Qwen-vs-GLM rivalry has settled into roles: Qwen wins on breadth of the family (sizes, modalities, quantization-friendly small models), GLM wins on flagship agentic capability. GLM-5.3’s security and agentic numbers have no Qwen equivalent this cycle, and Qwen’s flagship has not led Terminal-Bench. If your stack needs a 32B model on a single GPU, Qwen is still the conversation. If you need a frontier open-weights agent, GLM-5.3 is now the default.

The last comparative point is the least fashionable and the most useful: nobody serious deploys exactly one model in 2026. The practical pattern I see in production stacks is a router in front of three or four models — a cheap high-throughput model for extraction and classification, a mid-tier model for routine codegen, and a frontier model like GLM-5.3 or Fable 5 for the long-horizon, high-stakes tasks. Read the GLM-5.3 comparisons above as inputs to that router, not as a winner-take-all verdict. The right question is not “which model is best” but “which model earns which share of the traffic,” and GLM-5.3’s share is the agentic, long-horizon, security-adjacent slice that used to be closed-model-only territory.

The one comparison I refuse to make is a leaderboard total. Aggregating these suites into a single “who wins” number is how marketing departments eat benchmark reports, and it is not how you should buy a model. GLM-5.3 wins the agentic bucket outright. DeepSeek wins the cost-per-token bucket. Fable 5 wins the judgment bucket. Qwen wins the family-breadth bucket. Pick your bucket, then your model — that is the entire decision framework, and it has not changed in two years.

Cost and access

Let me be precise about pricing, because there is a lot of sloppy writing headed your way. GLM-5.3’s API price is not public. At launch, the API is staged behind the same safety review as the weights, and Zhipu has not published 5.3 rates. The reference point is GLM-5.2, which shipped at $1.40 per 1M input tokens and $4.40 per 1M output tokens, and the expectation is that 5.3 lands in the same band. Anyone quoting you a GLM-5.3 price today is inventing it.

What is public: GLM Coding Plan and ZCode access, with all user quotas reset on launch day. That is how Zhipu has been distributing flagship capability to heavy users — a subscription wrapper around the model, with the API for programmatic access. For self-host, GLM-5.2’s MIT weights remain the path until the staged 5.3 weights clear the review, and 5.2 is a completely legitimate fallback rather than a consolation prize.

What a heavy agent session costs (GLM-5.2 rates) Cost of a 1M-token agent session (GLM-5.2 rates) $1.40 / 1M input · $4.40 / 1M output · 5.3 expected in the same band input tokens output tokens $0 $0.50 $1.00 $1.50 $2.00 1M in + 16K out $1.47 1M in + 64K out $1.68 1M in + 128K out $1.96
Derived from GLM-5.2's published $1.40/$4.40 per 1M tokens, not a 5.3 price. Note how little of the bar is output cost at 16K outputs, and how much it grows at the 128K output ceiling — long-horizon agents are output-token events.

The cost math is worth internalizing because it changes what “expensive” means for agent workloads. A 1M-token-context agent session burns its context every few turns — the model re-reads the repo, the tool results, the plan — and the output tokens from planning and recovery are where the spend compounds. At GLM-5.2 rates, a session with 1M input tokens and 128K output tokens costs about $1.96. Run that across a fleet of agents doing eight-hour shifts and the bill is real, which is exactly why the GLM-5.3 post-training efficiency matters: fewer retries per task is the difference between a $2 session and a $6 session on the same job.

The access paths in order of friction: (1) GLM Coding Plan and ZCode, quotas reset at launch, fastest way to test 5.3 today; (2) the API once it clears review, expected in the same price band as 5.2; (3) self-host on staged weights, roughly two weeks out, running 5.2 under MIT in the interim. If your compliance regime requires you to run weights you can audit, you will be on 5.2 for the next two weeks and on 5.3 shortly after — and the migration risk is low, because the token formats and context behavior are unchanged.

When you compare GLM-5.3’s expected pricing against the closed frontier, remember the license premium baked into that comparison. At GLM-5.2’s rates, a 1M-in/128K-out session is under $2; the equivalent session on a closed frontier model is typically several times that before volume discounts. The open-weights price advantage is not a rounding error — it is the whole reason the self-host path matters, because at agentic workload volumes the difference between $2 and $8 per session is the difference between a pilot and a fleet. The 5.3 numbers that matter for cost are therefore not the per-token rates (expected in 5.2’s band) but the post-training efficiency that cuts retries per task — a per-task cost reduction that no pricing sheet will show you.

How to use GLM-5.3

The practical part. GLM-5.3 exposes a standard OpenAI-compatible chat-completions surface, which means it drops into the tooling you already have — the same client libraries, the same function-calling schema, the same evals harness. The model string and endpoint are staged with the safety review, so verify the current identifiers against Z.ai’s docs before you wire this in, but the shape below is what you are building toward.

from openai import OpenAI
client = OpenAI(
api_key="your-zai-api-key", # from the Z.ai console
base_url="https://api.z.ai/api/paas/v4", # verify current endpoint at launch docs
)
resp = client.chat.completions.create(
model="glm-5.3", # exact model string staged; check docs
messages=[
{
"role": "system",
"content": (
"You are GLM-5.3. You plan before you edit, you write the "
"failing test first, and you ask for help when a change "
"touches code you do not understand."
),
},
{
"role": "user",
"content": (
"In this repo, refactor the billing service to the repository "
"pattern, add a failing test for the proration edge case, make "
"it pass, and update the docs. Do not touch the auth module."
),
},
],
max_tokens=8192,
temperature=0.2,
)
print(resp.choices[0].message.content)

If you are wiring this into an agent harness rather than a chat, the parts that matter are the ones the benchmarks reward. Set a long horizon and a hard time budget; GLM-5.3 is built for eight-hour runs and the ExploitGym numbers show it sustains effort. Give it read access to the repo before you ask for a plan — the model’s biggest failure-mode improvement is planning before editing, and it needs the files to do that. Structure recovery into the loop: when a tool call fails, return the error text to the model rather than swallowing it, because error-driven recovery is where the Terminal-Bench gains came from. And cap temperature low (0.1-0.3) for code edits; the agentic post-training does the exploration, you do not need sampling heat on top of it.

The self-host path matters for two audiences: teams that need auditability and teams that need to avoid API egress for code. Until the staged weights land, GLM-5.2 is the MIT-licensed fallback and it is a good one. When 5.3’s weights clear, expect the same serving stack that 5.2 used — the model is the same 743B MoE family, the sparse-attention IndexShare machinery is unchanged, and the quantization story that made 5.2 shippable on commodity clusters should carry straight across. If you already serve GLM-5.2, the 5.3 upgrade should be a model-file swap, not a re-architecture.

One deployment note I will add from experience: do not put a 1M-context model behind a prompt template that fills context with boilerplate. The 1M window is for code, not for marketing copy. Every token in the window is a token the model has to attend over, and the IndexShare savings do not mean the model thinks faster on irrelevant content. Keep the system prompt tight, keep the injected context to the repo and the task, and let the 1M window be the safety margin for big codebases rather than the default working set.

Finally, build the observability before you build the agent. The failure modes of an agentic model show up in the logs, not in the chat: watch the ratio of tool calls that succeed on the first try, the number of retries per task, the time to first edit, and the frequency of the model asking for clarification. GLM-5.3’s post-training should move all four in the same direction, and if it does not, your harness — not the model — is probably the problem. Instrument the loop, baseline it on GLM-5.2, then swap in 5.3 and read the delta. That is the closest thing to a free eval that exists.

Honest caveats

This section is the reason the review is worth reading, so I am going to be direct about what is not yet true about GLM-5.3.

The benchmarks are vendor-published. Every number above came from Zhipu’s launch materials, not from an independent harness I ran. The Terminal-Bench, DeepSWE, ALE, and ExploitGym figures are all credible and internally consistent, and they line up with the failure-mode changes I observed — but “line up with my observations” is not “reproduced by an independent lab.” The weights are staged precisely because Zhipu does not want everyone running their own evals yet. Re-run everything when they land. If the 28.3 collapses under third-party testing, this review’s verdict collapses with it, and I am comfortable with that.

The vendor claims are bigger than the numbers. “~50% better at coding” and “approaching Claude Fable 5” are not benchmark results; they are framing. The reproducible part of the coding claim is the benchmark convergence, which is strong. The non-reproducible part — “approaching Fable 5” — is a direction, not a position. Plan around the numbers you can verify; treat the positioning as marketing intent.

No vision, and the roadmap is unclear. The top community request (vision) was explicitly deferred, and there is no announced date for a GLM-5.3V. If your product needs multimodal, this release does not move you, and GLM-5V-Turbo is a sidecar, not a flagship. The June 29 poll from founder Jie Tang was the tell that Zhipu knows the demand — and shipped anyway, which tells you what it thinks the next frontier actually is.

The two-week weight window is a promise, not a contract. Safety reviews slip. If the review turns up something the staged numbers do not show, the window extends and the API gate stays shut. Budget for the possibility that “about two weeks” becomes “the fall,” and keep your GLM-5.2 deployment healthy — which you should be doing anyway, because it is the licensed fallback.

Open weights are staged, not free of friction. When 5.3’s weights do ship, they are staged behind a review process — meaning distribution may be gated by registration, use-case attestation, or export restrictions in a way MIT weights were not. “Open weights” and “no strings” are different things, and GLM-5.2’s MIT license spoiled the community into expecting both. Read the actual license before you plan a product around the weights.

The cyber-defense numbers are the reason to be careful, not the reason to be afraid. The 54.4% ExploitBench and 105/130 ExploitGym completions are capabilities that defensive teams will use and that offensive actors will try to misuse. The safety staging is real and appropriate, and it also means the model you test today is not necessarily the model you deploy — the gated distribution may differ from the benchmarked weights. Treat both the capability and the gate as part of the product you are evaluating.

FAQ

Is GLM-5.3 open source?

Not day-one. The weights are staged behind a safety review, expected within about two weeks of the August 14, 2026 launch — a deliberate departure from GLM-5.2’s MIT-weights-on-day-one pattern. Until they land, GLM-5.2 remains the self-host fallback and the #1 open-weights model on the Artificial Analysis Index.

How much does GLM-5.3 cost?

The API price was not public at launch; the API is staged with the safety review. The reference is GLM-5.2 at $1.40 per 1M input and $4.40 per 1M output tokens, and 5.3 is expected in the same band. GLM Coding Plan and ZCode subscribers got quota resets on launch day.

Is GLM-5.3 better than DeepSeek?

On the published coding and agentic benchmarks, yes: it is the first open-source model to lead Terminal-Bench 3.0 and its DeepSWE v1.1 score of 66.9 is the strongest open-weights number of the year. But DeepSeek still leads on cost and throughput at scale, and Zhipu’s numbers are not independently verified yet.

Does GLM-5.3 have vision?

No. It is text-only. Vision was the top community request (per founder Jie Tang’s June 29 poll) and was deferred. Multimodal GLM work still runs on GLM-5V-Turbo.

What is GLM-5.3’s context window?

1M-token lossless context with up to 128K output tokens, inherited from GLM-5.2, using sparse attention with IndexShare indexer reuse (2.9x FLOP reduction at 1M context).

How is GLM-5.3 different from GLM-5.2?

Same 743B MoE base, so the entire delta is post-training. Terminal-Bench 3.0 went 4.6 to 28.3, DeepSWE v1.1 went 46.2 to 66.9, ExploitBench went 24.4% to 54.4%, and ExploitGym completions roughly tripled.

Is GLM-5.3 good at coding?

It is the release’s core claim: roughly 50% better than 5.2 on internal evals and approaching Claude Fable 5 on coding and agentic benchmarks. The benchmark evidence supports the direction; treat the Fable 5 comparison as unverified until third parties test it.

When was GLM-5.3 released?

August 14, 2026, the fourth flagship release in about six months — GLM-5 on February 12, GLM-5.1 in early April, GLM-5.2 on June 16, GLM-5.3 on August 14.

Further reading

All benchmark figures are as published by Zhipu AI at the August 14, 2026 launch. GLM-5.3 pricing is not public at launch; cost math uses GLM-5.2’s published $1.40 in / $4.40 out per 1M tokens as the reference band. Re-verify benchmarks independently once the staged weights ship.

Frequently Asked Questions

Is GLM-5.3 open source?

Not day-one. GLM-5.3's weights are staged behind a safety review and are expected within about two weeks of the August 14, 2026 launch. That is a deliberate departure from GLM-5.2, which shipped MIT-licensed weights on launch day. If you need open weights today, GLM-5.2 remains the self-host fallback and the #1 open-weights model on the Artificial Analysis Index.

How much does GLM-5.3 cost?

GLM-5.3 API pricing was not fully public at launch because the API is staged behind the same safety review as the weights. The reference point is GLM-5.2, published at $1.40 per 1M input tokens and $4.40 per 1M output tokens; 5.3 is expected to land in the same band. GLM Coding Plan and ZCode subscribers had all quotas reset on launch day.

Is GLM-5.3 better than DeepSeek?

On coding and agentic benchmarks, Zhipu claims GLM-5.3 leads the open-weights field and surpasses other Chinese models on coding experience. On Terminal-Bench 3.0 it is the first open-source model to reach the top tier. The caveat: these are Zhipu-published numbers as of launch, independent verification is pending, and DeepSeek still wins on some cost and throughput metrics.

Does GLM-5.3 have vision?

No. GLM-5.3 is text-only. Vision was the top community request — founder Jie Tang ran a poll on June 29, 2026 — but it was not included in this launch. If you need a multimodal GLM, GLM-5V-Turbo from April 2026 remains the option.

What is GLM-5.3's context window?

1M-token lossless context, inherited from GLM-5.2, with up to 128K output tokens. The 1M-token lineage relies on sparse attention with IndexShare indexer reuse, which GLM-5.2 showed cuts FLOPs by 2.9x at 1M context.

How is GLM-5.3 different from GLM-5.2?

Same 743B-parameter MoE base — the architecture did not grow. All the gains came from post-training. Terminal-Bench 3.0 jumped from 4.6 to 28.3 (6x), DeepSWE v1.1 went from 46.2 to 66.9, ExploitBench went from 24.4% to 54.4%, and ExploitGym task completions roughly tripled.

Is GLM-5.3 good at coding?

That is its core pitch. Zhipu says coding is roughly 50% better than GLM-5.2 on internal evals and that the model is approaching Claude Fable 5 on coding and agentic benchmarks — which would put it at the frontier. Treat internal-eval claims with skepticism until third parties reproduce them.

When was GLM-5.3 released?

August 14, 2026. It is Zhipu's fourth flagship release in about six months: GLM-5 on February 12, GLM-5.1 in early April, GLM-5.2 on June 16, and GLM-5.3 on August 14.