Help Center/ MaaS/ Model List
Updated on 2026-09-21 GMT+08:00

Model List

MaaS provides multiple models for you to use. You can easily integrate model services into your business by following tutorials or API documentation.

Text Generation

Tutorial: Text Generation; API: Sending a Chat Request (Chat/Post); price: MaaS Text Generation Models.

Table 1 Model list

Model Parameter (Model ID)

Supported Capabilities

Length Limit

Default Rate Limit

glm-5.3

Deep thinking (cannot be disabled)

Function calling

Prefix caching

Context length: 1M

Max input length: 1M

Max output length: 128K

Max CoT length: 128K

  • TPM: 1,000,000
  • RPM: 100

glm-5.2

Deep thinking (configurable)

Function calling

Context length: 1M

Max input length: 1M

Max output length: 128K

Max CoT length: 64K

  • TPM: 1,000,000
  • RPM: 100

glm-5.1

Deep thinking (configurable)

Function calling

Context length: 198K

Max input length: 192K

Max output length: 128K

Max CoT length: 96K

  • TPM: 1,000,000
  • RPM: 100

deepseek-v4.1-flash

Deep thinking (configurable)

Function calling

Prefix caching

Context length: 1M

Max input length: 1M

Max output length: 384K

Max CoT length: 96K

  • TPM: 1,000,000
  • RPM:100

deepseek-v4-flash

Deep thinking (configurable)

Function calling

Prefix completion

Context length: 1M

Max input length: 1M

Max output length: 384K

Max CoT length: 96K

  • TPM: 60,000
  • RPM: 15

deepseek-v4-pro

Deep thinking (configurable)

Function calling

Prefix completion

Context length: 1M

Max input length: 1M

Max output length: 128K

Max CoT length: 96K

  • TPM: 30,000
  • RPM: 3

Deep Thinking

Tutorial: Deep Thinking; API: Sending a Chat Request (Chat/Post); price: MaaS Text Generation Models.

Table 2 Model list

Model Parameter (Model ID)

Supported Capabilities

Length Limit

Default Rate Limit

glm-5.3

Deep thinking (cannot be disabled)

Function calling

Prefix caching

Context length: 1M

Max input length: 1M

Max output length: 128K

Max CoT length: 128K

  • TPM: 1,000,000
  • RPM: 100

glm-5.2

Deep thinking (configurable)

Function calling

Context length: 1M

Max input length: 1M

Max output length: 128K

Max CoT length: 64K

  • TPM: 1,000,000
  • RPM: 100

glm-5.1

Deep thinking (configurable)

Function calling

Context length: 198K

Max input length: 192K

Max output length: 128K

Max CoT length: 96K

  • TPM: 1,000,000
  • RPM: 100

deepseek-v4.1-flash

Deep thinking (configurable)

Function calling

Prefix caching

Context length: 1M

Max input length: 1M

Max output length: 384K

Max CoT length: 96K

  • TPM: 1,000,000
  • RPM:100

deepseek-v4-flash

Deep thinking (configurable)

Function calling

Prefix completion

Context length: 1M

Max input length: 1M

Max output length: 384K

Max CoT length: 96K

  • TPM: 60,000
  • RPM: 15

deepseek-v4-pro

Deep thinking (configurable)

Function calling

Prefix completion

Context length: 1M

Max input length: 1M

Max output length: 128K

Max CoT length: 96K

  • TPM: 30,000
  • RPM: 3