Explore indexBack to Models
GLM-5.3-FlashX
A high-speed inference large language model by Zhipu AI built on a 320B total / 18B active parameter hybrid attention architecture with a 1M context window and generation speeds up to 200 tokens/s.
A high-speed inference large language model by Zhipu AI built on a 320B total / 18B active parameter hybrid attention architecture with a 1M context window and generation speeds up to 200 tokens/s.