Help Center/ MaaS/ Model Calling/ Recommended Models
Updated on 2026-07-08 GMT+08:00

Recommended Models

MaaS provides multiple models for you to use. You can easily integrate model services into your business by following tutorials or API documentation.

Recommended models

GLM-5.2

GLM-5.1

GLM-5

GLM-5.2 is Zhipu AI's next-generation flagship model, officially launched and open-sourced on June 17, 2026. Its core features include:

Enhanced long-horizon task capability: reduces context drift and goal forgetting in complex tasks.

State-of-the-art coding and long-horizon task evaluation: achieves open-source SOTA performance, delivering greater stability in complex system engineering and deep debugging.

Improved real-world development experience: more reliable project-level context handling, adherence to engineering standards, and multi-platform development.

Support for 198K sequence length, with maximum input of 192K, maximum output of 128K, and maximum thought chain length of 64K.

Support for advanced functions, including thought chains, function calling, and prefix continuation.

GLM-5.1, the latest flagship model from Zhipu, delivers enhanced coding capabilities and superior performance on long-horizon tasks. It can autonomously operate for up to eight hours on a single task, covering the full cycle from planning and execution to iterative optimization, and producing engineering-level results.

While GLM-5.1 matches Claude Opus 4.6 in general intelligence and raw coding proficiency, it significantly outperforms the global frontier in long-horizon sustained execution. It excels at autonomous, multi-stage tasks over extended periods, making it the ideal foundation for building highly resilient autonomous agents and long-horizon coding engines.

It supports sequence lengths up to 198K, with a maximum input of 192K, a maximum output of 128K, and a maximum thought chain of 96K.

It supports chain-of-thought, function calling, and prefix continuation.

Compared with GLM-4.5, GLM-5 expands total parameters from 355B to 744B (boosting active parameters from 32B to 40B) and increases pre-training data from 23T to 28.5T tokens. In addition, GLM-5 integrates the DeepSeek sparse attention (DSA) mechanism, which significantly reduces deployment costs while maintaining the long context capability.

Compared to GLM-4.7, GLM-5 delivers substantial performance gains across a wide spectrum of academic benchmarks. It establishes a new state-of-the-art (SOTA) among global open-source models in reasoning, coding, and agentic tasks, further closing the capability gap with frontier closed-source models.

It supports sequence lengths up to 198K, with a maximum input of 192K, a maximum output of 64K, and a maximum thought chain of 64K.

It supports chain-of-thought, function calling, and prefix continuation.

Deep Thinking

For tutorials, see Deep Thinking. For API documentation, see Sending a Chat Request (Chat/Post).

Table 1 Recommended models

Model Name

Version

model Parameter

Capability

Context Length

Max Input Length

Max Output Length

Max Chain-of-Thought Length

Supported Region

GLM-5.2

20260617

glm-5.2

Deep thinking (configurable)

Function calling

Prefix continuation

198K

192K

128K

64K

CN-Hong Kong

GLM-5.1

20260407

glm-5.1

Deep thinking (configurable)

Function calling

Prefix continuation

198K

192K

128K

96K

CN-Hong Kong

GLM-5

20260212

glm-5

Deep thinking (configurable)

Function calling

Prefix continuation

198K

192K

64K

64K

CN-Hong Kong

DeepSeek-V4-Pro

20260424

deepseek-v4-pro

Deep thinking (configurable)

Function calling

Prefix continuation

1M

1M

128K

96K

CN-Hong Kong

DeepSeek-V4-Flash

20260424

deepseek-v4-flash

Deep thinking (configurable)

Function calling

Prefix continuation

1M

1M

128K

96K

CN-Hong Kong

DeepSeek-V3.2

20251215

deepseek-v3.2

Deep thinking (configurable)

Function calling

Prefix continuation

160K

128K

32K

32K

CN-Hong Kong

DeepSeek-V3.1

20251124

deepseek-v3.1-terminus

Deep thinking (configurable)

Function calling

Prefix continuation

128K

96K

32K

32K

CN-Hong Kong

Text Generation

For tutorials, see Text Generation Overview. For API documentation, see Sending a Chat Request (Chat/Post).

Table 2 Recommended models

Model Name

Version

model Parameter

Capability

Context Length

Max Input Length

Max Output Length

Max Chain-of-Thought Length

Supported Region

GLM-5.2

20260617

glm-5.2

Deep thinking (configurable)

Function calling

Prefix continuation

198K

192K

128K

64K

CN-Hong Kong

GLM-5.1

20260407

glm-5.1

Deep thinking (configurable)

Function calling

Prefix continuation

198K

192K

128K

96K

CN-Hong Kong

GLM-5

20260212

glm-5

Deep thinking (configurable)

Function calling

Prefix continuation

198K

192K

64K

64K

CN-Hong Kong

DeepSeek-V4-Pro

20260424

deepseek-v4-pro

Deep thinking (configurable)

Function calling

Prefix continuation

1M

1M

128K

96K

CN-Hong Kong

DeepSeek-V4-Flash

20260424

deepseek-v4-flash

Deep thinking (configurable)

Function calling

Prefix continuation

1M

1M

128K

96K

CN-Hong Kong

DeepSeek-V3

20250929

DeepSeek-V3

Function calling

128K

128K

32K

N/A

CN-Hong Kong

DeepSeek-R1-0528

20250929

deepseek-r1-250528

Deep thinking

Function calling

Prefix continuation

128K

96K

32K

32K

CN-Hong Kong