Umans GLM 5.3
115.6tok/s
throughput · p50 · last 5 min
422ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h
GLM 5.3 is Z.ai's flagship open-weights model for coding and long-horizon agentic work: a 744B mixture-of-experts with 40B active parameters and a 1M-token context window, the successor to GLM 5.2 on the same base with large post-training gains on complex coding tasks. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off).
Trends
Speed over the last 90 days
now 108.3 tok/s
90 days agopre-release before Aug 29, 2026today
best 1.06s · Aug 29now 1.16s
90 days agopre-release before Aug 29, 2026today
Changelog
Events for Umans GLM 5.3
Aug 292026
Released pay-per-token: Umans GLM 5.3 Released
umans-glm-5.3 joins the lineup: Z.ai's flagship open-weights model for coding and long-horizon agentic work - a 744B mixture-of-experts with 40B active parameters, the successor to GLM 5.2 on the same base with large post-training gains, on a 1M context window. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off). Billed per token: $1.40 / $4.40 / $0.26 per 1M (input / output / cache read). It succeeds GLM 5.2, which is deprecated.