As an AI consultancy from Rosenheim, we support companies in Germany, Austria and Switzerland with the GDPR-compliant use of open-weight models. For Z.ai’s (Zhipu AI) GLM family, the decisive question is which version is actually available as open weights – and as of 17 August 2026 that is not the newest one.
Status August 2026 at a glance
- GLM-5.3 (14 Aug 2026) is the current flagship, but so far usable only via the Z.ai API and the GLM Coding Plan. Z.ai held the weights back and announced their release for roughly two weeks after launch.
- GLM-5.2 (June 2026) is therefore the most recent version with available weights under the MIT licence – and the practical choice for sovereign self-hosting.
- For image and video input there are dedicated models: GLM-5V-Turbo, GLM-4.6V and GLM-OCR. The GLM-5.x flagships themselves are text-only according to the Z.ai documentation.
GLM-5.3 – new post-training on a familiar base
GLM-5.3 was released on 14 August 2026. The way it got there is notable: Z.ai did not retrain the GLM-5.2 base model (approx. 744B parameters, Mixture-of-Experts). According to the vendor, all gains come from scaled post-training.
The jumps Z.ai reports mainly concern agentic tasks:
| Benchmark | GLM-5.3 | GLM-5.2 |
|---|---|---|
| Terminal-Bench 3.0 | 28.3 | 4.6 |
| DeepSWE v1.1 | 66.9 | 46.2 |
| Agents’ Last Exam (CLI) | 28.5 | 23.8 |
| CyberGym | 84.5% | 77.2% |
| ExploitBench | 54.4% | 24.4% |
Important context: these are vendor figures. Independent measurements – such as the Artificial Analysis Intelligence Index – are not yet available for GLM-5.3, because the weights are missing and the general API is not open yet. The Terminal-Bench 3.0 numbers are also not comparable to the older Terminal-Bench 2.x figures; the benchmark generation is considerably harder.
Technical specs (per docs.z.ai)
- 1M token context window, up to 128k token output
- Three reasoning effort levels:
low,highandmax–maxis the default and is recommended for complex coding - Reasoning can no longer be switched off: integrations using
thinking.type: "disabled"must move toenabledwithreasoning_effort: "low"before upgrading - Function calling, context caching, structured output, streaming
- Text input only for now
- Protocols: OpenAI Chat Completion, with OpenAI Response and Anthropic Messages announced
Why the weights are being held back
Z.ai postponed the release of the GLM-5.3 weights by roughly two weeks to complete a safety evaluation. The background is the model’s sharply increased ability to find and exploit software vulnerabilities. Z.ai reports 2,436 findings across 269 open-source projects, 53 of which were disclosed at launch and assigned CVEs.
Two things follow for companies. First, a self-hosting date for GLM-5.3 cannot yet be planned reliably. Second, the licence of the forthcoming weights is not yet confirmed – GLM-5, GLM-5.1 and GLM-5.2 shipped under MIT, so continuity is plausible but not a commitment. What ultimately counts is the LICENSE file in the repository.
GLM-5.2 – the basis for sovereign self-hosting
GLM-5.2 (June 2026) is available under the MIT licence with no regional usage restrictions – which is why it is currently the more interesting version from a compliance perspective:
- Free commercial use without licensing fees
- Modification, fine-tuning and deployment on your own infrastructure
- Full control over data and model, no vendor lock-in
Architecture
744B parameters with 40B active per token in a Mixture-of-Experts architecture. Z.ai highlights IndexShare: one indexer is reused across every four sparse attention layers, cutting per-token compute by 2.9x at a 1M token context. Two reasoning modes are available: High (fast) and Max (deeper for complex coding).
Independent assessment
On the Artificial Analysis Intelligence Index v4.1.1, GLM-5.2 scores 53, placing it among the strongest models with open weights. At its June launch the score was 51 on index v4.1, ahead of MiniMax-M3 (44), DeepSeek V4 Pro (44) and Kimi K2.6 (43).
This is where vendor and third-party measurements visibly diverge: Artificial Analysis flags GLM-5.2 as notably token-intensive – around 43,000 output tokens per task versus 26,000 for GLM-5.1 and 24,000 for MiniMax-M3. The raw intelligence ranking therefore says little about actual running costs. Anyone buying GLM-5.2 through an API should measure consumption against their own workload rather than calculating from the list price. Output speed was measured at around 152 tokens per second.
For context: other open models also moved forward over the summer of 2026. Kimi K3 from Moonshot AI achieves a higher score on the Artificial Analysis index, but at the time of our research its weights were announced and not yet published.
Pricing and availability
Z.ai API (official price list, USD per 1M tokens)
| Model | Input | Output |
|---|---|---|
| GLM-5.2 | 1.40 | 4.40 |
| GLM-5.1 | 1.40 | 4.40 |
| GLM-5-Turbo | 1.20 | 4.00 |
| GLM-5 | 1.00 | 3.20 |
| GLM-4.7 | 0.60 | 2.20 |
| GLM-4.7-FlashX | 0.07 | 0.40 |
| GLM-4.7-Flash | free | free |
| GLM-5V-Turbo (vision) | 1.20 | 4.00 |
| GLM-4.6V (vision) | 0.30 | 0.90 |
| GLM-OCR | 0.03 | 0.03 |
For GLM-5.2, Z.ai additionally lists a cache-hit price of USD 0.26 per 1M tokens. For GLM-5.3 no per-token price has been published; the general API is marked “coming soon” in the documentation.
GLM Coding Plan
Z.ai also sells coding access as a subscription with a quota system (Lite, Pro, Max plus a Team variant), usable in coding agents such as ZCode, Claude Code, Cline or Roo Code. The entry-level Lite tier is around USD 18 per month; figures for the higher tiers vary by source, which is why we do not quote them here – please check directly with Z.ai. GLM-5.3 is currently accessible only through this route and via ZCode.
Third-party providers
- OpenRouter carries the GLM family broadly; GLM-5.2 is listed there at roughly USD 0.49 input and USD 1.54 output per 1M tokens (1.05M context), plus a free 128k variant. GLM-5.3 was not listed there at the time of our research.
- Together, Fireworks, DeepInfra and other inference providers carry the open GLM models – predominantly on US infrastructure.
- Hyperscaler catalogues (AWS Bedrock, Azure AI Foundry, Google Vertex AI) do not offer GLM as a first-party model; older GLM releases are partly available through marketplace catalogues such as NVIDIA NIM.
GDPR assessment
This is a factual assessment, not legal advice. The specific evaluation depends on your deployment scenario and belongs with your data protection function.
Who actually processes the data?
It is important to distinguish two offerings:
- z.ai is the international offering. Per its privacy policy, the operator is JINGSHENG HENGXING TECHNOLOGY PTE. LTD., registered in Singapore. Processing generally takes place out of Singapore. For the API services, Z.ai states that prompts and outputs are not stored but processed in real time.
- open.bigmodel.cn is the Chinese platform of the Zhipu group.
What follows from this
There is no European Commission adequacy decision under Art. 45 GDPR for either China or Singapore. Transferring personal data to these services therefore constitutes a third-country transfer under Art. 44 et seq. GDPR and requires appropriate safeguards under Art. 46 GDPR – in practice standard contractual clauses supplemented by a transfer impact assessment that also evaluates government access possibilities in the destination country. A data processing agreement under Art. 28 GDPR is required in addition.
The commitment not to store API content reduces risk but replaces neither the legal basis for the transfer nor its documentation. For personal or business-critical data we therefore recommend restricting direct API use to non-personal evaluation and prototyping scenarios.
The clean route: EU operation
Because GLM-5.2 and GLM-5.1 are available under the MIT licence, the third-country question can be solved constructively – the model moves to the data instead of the other way round:
- Self-hosting on your own or rented GPU infrastructure in an EU data centre
- European inference providers: Scaleway has carried GLM-5.2 in its Generative APIs since 26 June 2026 and states that inference runs on European infrastructure with no routing to the model vendor. Nebius also offers GLM models on European infrastructure. In both cases, the data processing agreement, sub-processor list and storage locations should be clarified directly with the provider.
- Private cloud or on-premise for scenarios involving particularly sensitive data
This is exactly the pattern we implement with CompanyGPT: open models in your own EU environment, with access control, logging and tenant separation.
Hardware requirements for self-hosting
GLM-5.2 is a very large model. Even though only around 40B parameters are active per token, all expert weights must reside in VRAM. As a rough guide for the weights alone:
- BF16/FP16: around 1,500 GB VRAM – in practice roughly 16 GPUs of the H100/H200 class
- FP8/INT8: around 750 GB VRAM – an 8-GPU H200 node is the common reference and leaves headroom for the KV cache
- INT4: around 370 GB VRAM
The KV cache comes on top and is substantial at a 1M token context. vLLM and SGLang are the established serving stacks; pre-quantised checkpoints are available. If your utilisation is not continuous, a European inference provider is usually more economical than owning hardware.
Superseded models
- GLM-5 (February 2026) and GLM-5.1 (April 2026) have been superseded by GLM-5.2. Both remain available as open weights and are still in the Z.ai price list.
- GLM-5-Turbo (March 2026) has been overtaken by the 5.1/5.2 generations.
- Older releases such as GLM-4.5, GLM-4.6 and GLM-4.7 are still offered, partly as free Flash variants, but are no longer the first choice for new projects.
- In third-party catalogues (e.g. NVIDIA NIM), individual GLM releases have already been deprecated. For existing integrations, check your own provider’s announcements, not just Z.ai’s.
Our recommendation
For production use involving personal data, the relevant version right now is GLM-5.2 – not GLM-5.3. Only GLM-5.2 can be run with open weights under the MIT licence in an EU data centre and thus embedded in a clean data protection framework.
GLM-5.3 is interesting for evaluations in a coding context, as long as no personal or confidential data is involved. Once the weights and the licence are published, a fresh assessment is worthwhile – then with independent benchmark data rather than vendor figures alone.
We support you with:
- Selecting and evaluating open models for your specific use case
- Infrastructure planning, hardware sizing and cost modelling
- Deployment on EU infrastructure or connection to a European inference provider
- Data protection and AI Act documentation including transfer impact assessment
- Fine-tuning and integration into existing systems
With a 1M token context window, an MIT licence and strong agentic capabilities, GLM-5.2 offers a solid foundation for sovereign AI applications – provided the infrastructure question is settled. That is exactly where we help with CompanyGPT.
