Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM Zhipu AI China

GLM-5.3

GLM-5.3 and GLM-5.2 by Z.ai (Zhipu AI): release, benchmarks, pricing and open weights explained. AI consultancy from Rosenheim on GDPR-compliant GLM self-hosting in the EU.

License MIT (GLM-5, GLM-5.1, GLM-5.2); GLM-5.3 licence still open
GDPR Hosting Available
Context 1M (GLM-5.3 and GLM-5.2), 200k (GLM-5V-Turbo, GLM-5.1) Tokens
Modality Text → Text

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
GLM-5.3
14 August 2026
Same base model as GLM-5.2 (approx. 744B parameters, MoE) – Z.ai states all gains come from scaled post-training alone 1M token context window, up to 128k token output (source: docs.z.ai) Three reasoning effort levels: low, high and max (max is the default and recommended for complex coding) Clear jumps on agentic benchmarks: Terminal-Bench 3.0 28.3 (GLM-5.2: 4.6), DeepSWE v1.1 66.9 (46.2), Agents' Last Exam CLI 28.5 (23.8) Function calling, context caching and structured output; compatible with OpenAI Chat Completion, OpenAI Response and Anthropic Messages protocols Pronounced security capabilities: CyberGym 84.5% (77.2%), ExploitBench 54.4% (24.4%)
Weights not published as of 17 Aug 2026 – self-hosting currently not possible Licence of the forthcoming weights not yet confirmed by Z.ai (GLM-5.1 and GLM-5.2 shipped under MIT) No token price list, no EU hosting offering, no independent measurements (e.g. Artificial Analysis) available Reasoning can no longer be switched off – existing integrations using thinking.type: disabled must be adapted The very security capabilities Z.ai lists as a strength are the reason for the delayed weights release (dual-use risk)
Current
GLM-5.2 Recommended
June 2026
Latest GLM model with actually available weights – the practical basis for sovereign self-hosting Open source under MIT licence, no regional usage restrictions 1M token context window, up to 131k token output Artificial Analysis Intelligence Index v4.1.1: 53 – among the highest scores for models with open weights 744B parameters (40B active) MoE architecture with IndexShare (2.9x less compute per token at max context) Two reasoning modes: High (fast) and Max (deeper for complex coding) Usable via European providers such as Scaleway without your own GPU infrastructure
Token-intensive: Artificial Analysis measures around 43k output tokens per task (GLM-5.1: 26k) – verbosity drives up cost No native EU offering from Z.ai itself; GDPR compliance requires self-hosting or an EU provider High hardware requirements (1M context especially memory-intensive) Text-only – image and video input only via the separate GLM-5V / GLM-4.6V models
Current
GLM-5V-Turbo
April 2026
Natively multimodal: image, video and text input 200k token context window, up to 131k token output Designed for agentic engineering workflows and GUI operation
No known open weights – self-hosting therefore not possible No EU hosting offering
Current
GLM-5.1
April 2026
Long-horizon agentic tasks across thousands of tool calls Open source under MIT licence (unrestricted commercial use) 200k token context window, 131k token output 744B parameters (40B active) MoE architecture More economical in output than GLM-5.2 (approx. 26k instead of 43k tokens per task per Artificial Analysis)
Surpassed by GLM-5.2 in coding, context length and agentic tasks No native EU offering from Z.ai itself Requires own infrastructure or an EU provider for GDPR compliance
Deprecated
GLM-5
February 2026
744B parameters (40B active) with MoE architecture 200k token context window Open source (MIT licence) Strong coding and reasoning performance Agentic AI capabilities
Surpassed by GLM-5.1 and GLM-5.2 in coding and agentic tasks No native EU cloud integration Requires own infrastructure for GDPR compliance
Deprecated

Use Cases

Typical applications for this model

Agentic AI workflows
Software engineering & coding
Research & science
Document analysis
Complex reasoning tasks
Long-form content creation
Multi-step task planning

Technical Details

API, features and capabilities

API & Availability
Availability Public
Throughput ~152 (GLM-5.2, measured by Artificial Analysis) Tokens/Sec
Features & Capabilities
Tool Use Function Calling Structured Output Reasoning Mode Web Browsing File Upload
Training & Knowledge
Knowledge Cutoff Late 2025
Fine-Tuning Available (Full Fine-tuning, LoRA)
Language Support
Best Quality English, Chinese
Supported Multilingual
Best quality in English and Chinese

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-Hosted
EU (your choice)
GLM-5.2 and GLM-5.1 are available as open weights under the MIT licence and can run on your own infrastructure in EU data centres. GLM-5.3 weights are not published yet.
Scaleway Generative APIs
EU (France)
GLM-5.2 available since 26 June 2026; per the provider, inference runs on Scaleway infrastructure in Europe with no routing to the model vendor. Verify DPA and terms directly with Scaleway.
Nebius
EU
GLM models on European infrastructure; GLM-5.2 supported from day zero per the provider. Verify DPA and terms directly with Nebius.
License & Hosting
License MIT (GLM-5, GLM-5.1, GLM-5.2); GLM-5.3 licence still open
Security Filters Configurable
On-Premise

Benchmarks

Performance comparison with standardized tests

Terminal-Bench 3.0 (GLM-5.3)
28.3 2026-08
DeepSWE v1.1 (GLM-5.3)
66.9 2026-08
Agents' Last Exam CLI (GLM-5.3)
28.5 2026-08
CyberGym (GLM-5.3)
84.5 2026-08
ExploitBench (GLM-5.3)
54.4 2026-08
Intelligence Index v4.1.1 (GLM-5.2)
53 2026-08
Terminal-Bench 3.0 (GLM-5.2)
4.6 2026-08
DeepSWE v1.1 (GLM-5.2)
46.2 2026-08
SWE-Bench Pro (GLM-5.2)
62.1 2026-06
Terminal-Bench 2.1 (GLM-5.2)
81 2026-06
SWE-Bench Pro (GLM-5.1)
58.4 2026-04
Terminal-Bench 2.0 (GLM-5.1)
63.5 2026-04
BrowseComp (GLM-5.1)
68 2026-04
MMLU
85 2026-02
SWE-bench Verified
77.8 2026-02
AIME 2025
84 2026-02
GSM8k
97 2026-02
GPQA
68.2 2026-02

As an AI consultancy from Rosenheim, we support companies in Germany, Austria and Switzerland with the GDPR-compliant use of open-weight models. For Z.ai’s (Zhipu AI) GLM family, the decisive question is which version is actually available as open weights – and as of 17 August 2026 that is not the newest one.

Status August 2026 at a glance

  • GLM-5.3 (14 Aug 2026) is the current flagship, but so far usable only via the Z.ai API and the GLM Coding Plan. Z.ai held the weights back and announced their release for roughly two weeks after launch.
  • GLM-5.2 (June 2026) is therefore the most recent version with available weights under the MIT licence – and the practical choice for sovereign self-hosting.
  • For image and video input there are dedicated models: GLM-5V-Turbo, GLM-4.6V and GLM-OCR. The GLM-5.x flagships themselves are text-only according to the Z.ai documentation.

GLM-5.3 – new post-training on a familiar base

GLM-5.3 was released on 14 August 2026. The way it got there is notable: Z.ai did not retrain the GLM-5.2 base model (approx. 744B parameters, Mixture-of-Experts). According to the vendor, all gains come from scaled post-training.

The jumps Z.ai reports mainly concern agentic tasks:

BenchmarkGLM-5.3GLM-5.2
Terminal-Bench 3.028.34.6
DeepSWE v1.166.946.2
Agents’ Last Exam (CLI)28.523.8
CyberGym84.5%77.2%
ExploitBench54.4%24.4%

Important context: these are vendor figures. Independent measurements – such as the Artificial Analysis Intelligence Index – are not yet available for GLM-5.3, because the weights are missing and the general API is not open yet. The Terminal-Bench 3.0 numbers are also not comparable to the older Terminal-Bench 2.x figures; the benchmark generation is considerably harder.

Technical specs (per docs.z.ai)

  • 1M token context window, up to 128k token output
  • Three reasoning effort levels: low, high and maxmax is the default and is recommended for complex coding
  • Reasoning can no longer be switched off: integrations using thinking.type: "disabled" must move to enabled with reasoning_effort: "low" before upgrading
  • Function calling, context caching, structured output, streaming
  • Text input only for now
  • Protocols: OpenAI Chat Completion, with OpenAI Response and Anthropic Messages announced

Why the weights are being held back

Z.ai postponed the release of the GLM-5.3 weights by roughly two weeks to complete a safety evaluation. The background is the model’s sharply increased ability to find and exploit software vulnerabilities. Z.ai reports 2,436 findings across 269 open-source projects, 53 of which were disclosed at launch and assigned CVEs.

Two things follow for companies. First, a self-hosting date for GLM-5.3 cannot yet be planned reliably. Second, the licence of the forthcoming weights is not yet confirmed – GLM-5, GLM-5.1 and GLM-5.2 shipped under MIT, so continuity is plausible but not a commitment. What ultimately counts is the LICENSE file in the repository.

GLM-5.2 – the basis for sovereign self-hosting

GLM-5.2 (June 2026) is available under the MIT licence with no regional usage restrictions – which is why it is currently the more interesting version from a compliance perspective:

  • Free commercial use without licensing fees
  • Modification, fine-tuning and deployment on your own infrastructure
  • Full control over data and model, no vendor lock-in

Architecture

744B parameters with 40B active per token in a Mixture-of-Experts architecture. Z.ai highlights IndexShare: one indexer is reused across every four sparse attention layers, cutting per-token compute by 2.9x at a 1M token context. Two reasoning modes are available: High (fast) and Max (deeper for complex coding).

Independent assessment

On the Artificial Analysis Intelligence Index v4.1.1, GLM-5.2 scores 53, placing it among the strongest models with open weights. At its June launch the score was 51 on index v4.1, ahead of MiniMax-M3 (44), DeepSeek V4 Pro (44) and Kimi K2.6 (43).

This is where vendor and third-party measurements visibly diverge: Artificial Analysis flags GLM-5.2 as notably token-intensive – around 43,000 output tokens per task versus 26,000 for GLM-5.1 and 24,000 for MiniMax-M3. The raw intelligence ranking therefore says little about actual running costs. Anyone buying GLM-5.2 through an API should measure consumption against their own workload rather than calculating from the list price. Output speed was measured at around 152 tokens per second.

For context: other open models also moved forward over the summer of 2026. Kimi K3 from Moonshot AI achieves a higher score on the Artificial Analysis index, but at the time of our research its weights were announced and not yet published.

Pricing and availability

Z.ai API (official price list, USD per 1M tokens)

ModelInputOutput
GLM-5.21.404.40
GLM-5.11.404.40
GLM-5-Turbo1.204.00
GLM-51.003.20
GLM-4.70.602.20
GLM-4.7-FlashX0.070.40
GLM-4.7-Flashfreefree
GLM-5V-Turbo (vision)1.204.00
GLM-4.6V (vision)0.300.90
GLM-OCR0.030.03

For GLM-5.2, Z.ai additionally lists a cache-hit price of USD 0.26 per 1M tokens. For GLM-5.3 no per-token price has been published; the general API is marked “coming soon” in the documentation.

GLM Coding Plan

Z.ai also sells coding access as a subscription with a quota system (Lite, Pro, Max plus a Team variant), usable in coding agents such as ZCode, Claude Code, Cline or Roo Code. The entry-level Lite tier is around USD 18 per month; figures for the higher tiers vary by source, which is why we do not quote them here – please check directly with Z.ai. GLM-5.3 is currently accessible only through this route and via ZCode.

Third-party providers

  • OpenRouter carries the GLM family broadly; GLM-5.2 is listed there at roughly USD 0.49 input and USD 1.54 output per 1M tokens (1.05M context), plus a free 128k variant. GLM-5.3 was not listed there at the time of our research.
  • Together, Fireworks, DeepInfra and other inference providers carry the open GLM models – predominantly on US infrastructure.
  • Hyperscaler catalogues (AWS Bedrock, Azure AI Foundry, Google Vertex AI) do not offer GLM as a first-party model; older GLM releases are partly available through marketplace catalogues such as NVIDIA NIM.

GDPR assessment

This is a factual assessment, not legal advice. The specific evaluation depends on your deployment scenario and belongs with your data protection function.

Who actually processes the data?

It is important to distinguish two offerings:

  • z.ai is the international offering. Per its privacy policy, the operator is JINGSHENG HENGXING TECHNOLOGY PTE. LTD., registered in Singapore. Processing generally takes place out of Singapore. For the API services, Z.ai states that prompts and outputs are not stored but processed in real time.
  • open.bigmodel.cn is the Chinese platform of the Zhipu group.

What follows from this

There is no European Commission adequacy decision under Art. 45 GDPR for either China or Singapore. Transferring personal data to these services therefore constitutes a third-country transfer under Art. 44 et seq. GDPR and requires appropriate safeguards under Art. 46 GDPR – in practice standard contractual clauses supplemented by a transfer impact assessment that also evaluates government access possibilities in the destination country. A data processing agreement under Art. 28 GDPR is required in addition.

The commitment not to store API content reduces risk but replaces neither the legal basis for the transfer nor its documentation. For personal or business-critical data we therefore recommend restricting direct API use to non-personal evaluation and prototyping scenarios.

The clean route: EU operation

Because GLM-5.2 and GLM-5.1 are available under the MIT licence, the third-country question can be solved constructively – the model moves to the data instead of the other way round:

  • Self-hosting on your own or rented GPU infrastructure in an EU data centre
  • European inference providers: Scaleway has carried GLM-5.2 in its Generative APIs since 26 June 2026 and states that inference runs on European infrastructure with no routing to the model vendor. Nebius also offers GLM models on European infrastructure. In both cases, the data processing agreement, sub-processor list and storage locations should be clarified directly with the provider.
  • Private cloud or on-premise for scenarios involving particularly sensitive data

This is exactly the pattern we implement with CompanyGPT: open models in your own EU environment, with access control, logging and tenant separation.

Hardware requirements for self-hosting

GLM-5.2 is a very large model. Even though only around 40B parameters are active per token, all expert weights must reside in VRAM. As a rough guide for the weights alone:

  • BF16/FP16: around 1,500 GB VRAM – in practice roughly 16 GPUs of the H100/H200 class
  • FP8/INT8: around 750 GB VRAM – an 8-GPU H200 node is the common reference and leaves headroom for the KV cache
  • INT4: around 370 GB VRAM

The KV cache comes on top and is substantial at a 1M token context. vLLM and SGLang are the established serving stacks; pre-quantised checkpoints are available. If your utilisation is not continuous, a European inference provider is usually more economical than owning hardware.

Superseded models

  • GLM-5 (February 2026) and GLM-5.1 (April 2026) have been superseded by GLM-5.2. Both remain available as open weights and are still in the Z.ai price list.
  • GLM-5-Turbo (March 2026) has been overtaken by the 5.1/5.2 generations.
  • Older releases such as GLM-4.5, GLM-4.6 and GLM-4.7 are still offered, partly as free Flash variants, but are no longer the first choice for new projects.
  • In third-party catalogues (e.g. NVIDIA NIM), individual GLM releases have already been deprecated. For existing integrations, check your own provider’s announcements, not just Z.ai’s.

Our recommendation

For production use involving personal data, the relevant version right now is GLM-5.2 – not GLM-5.3. Only GLM-5.2 can be run with open weights under the MIT licence in an EU data centre and thus embedded in a clean data protection framework.

GLM-5.3 is interesting for evaluations in a coding context, as long as no personal or confidential data is involved. Once the weights and the licence are published, a fresh assessment is worthwhile – then with independent benchmark data rather than vendor figures alone.

We support you with:

  • Selecting and evaluating open models for your specific use case
  • Infrastructure planning, hardware sizing and cost modelling
  • Deployment on EU infrastructure or connection to a European inference provider
  • Data protection and AI Act documentation including transfer impact assessment
  • Fine-tuning and integration into existing systems

With a 1M token context window, an MIT licence and strong agentic capabilities, GLM-5.2 offers a solid foundation for sovereign AI applications – provided the infrastructure question is settled. That is exactly where we help with CompanyGPT.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Consultation for this model?

We help you select and integrate the right AI model for your use case.