umans/status/umans-glm-5.3
Live · refreshes every 30s
← all models
Umans GLM 5.3
umans-glm-5.3 · GLM-5.3 · Z.ai
Operational
115.6tok/s
throughput · p50 · last 5 min
422ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h

GLM 5.3 is Z.ai's flagship open-weights model for coding and long-horizon agentic work: a 744B mixture-of-experts with 40B active parameters and a 1M-token context window, the successor to GLM 5.2 on the same base with large post-training gains on complex coding tasks. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off).

90 days agoin production since Aug 29, 2026today
Context
1049K
Max output
131K
Recommended
131K
Vision
No
Tools
Yes
Reasoning
Always on
Weights
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
now 108.3 tok/s
90 days agopre-release before Aug 29, 2026today
TTFT p50 · time to first token, lower is better
best 1.06s · Aug 29now 1.16s
90 days agopre-release before Aug 29, 2026today
Changelog

Events for Umans GLM 5.3

incl. gateway-wide announcements
Aug 292026
Released pay-per-token: Umans GLM 5.3 Released
umans-glm-5.3 joins the lineup: Z.ai's flagship open-weights model for coding and long-horizon agentic work - a 744B mixture-of-experts with 40B active parameters, the successor to GLM 5.2 on the same base with large post-training gains, on a 1M context window. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off). Billed per token: $1.40 / $4.40 / $0.26 per 1M (input / output / cache read). It succeeds GLM 5.2, which is deprecated.