Z.ai/OtherActive

GLM 5

We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.

Context
198K
Max output
128K
Input
$0.60
Output
$1.92
DECISION SUMMARY

Recommended use cases

Strengths in this dataset

  • 198,000-token context window
  • text input
  • 19 supported API parameters listed

Limits and caveats

  • Provider behavior and pricing can change; verify the linked sources before production use.
CAPABILITIES

Capability

Model-native facts

Model
Reasoning
Supported
Open weights
Supported

Provider endpoint facts

Provider endpoint
Tool calling
Supported
Structured output
Supported
Streaming
Unknown
Prompt cache
Supported
Batch
Unknown
Fine-tuning
Unknown
PROVIDER PRICING

GLM 5 Provider pricing

Provider endpoint: z-ai/glm-5

Input
$0.60
per 1M tokens
Output
$1.92
per 1M tokens
Cached input
$0.12
per 1M tokens
Image output
Unknown
per 1M tokens
SOURCE RECORDS

Sources and verification

Hugging Face model card

Fields: tags, gated, license, summary, languages, library name, pipeline tag

OpenRouter Models API

Fields: identity, description, modalities, context window, maximum output, pricing, supported parameters

MODEL FAQ

Frequently asked questions

Answers are generated from the same sourced model and provider facts shown above.

We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.

Model specifications and prices may vary by provider and change over time. AIToolly displays sources and verification dates so users can confirm critical details before production use.