Zhipu launches GLM-5.3-Flash, its first natively multimodal model built for Chinese chips

2 days ago 1



China’s Z.ai, the company formerly known as Zhipu AI, released GLM-5.3-Flash on August 26, making it the first natively multimodal model in the GLM-5 series. The launch is notable for reasons beyond the model itself: it runs entirely on domestically produced AI accelerators, a pointed demonstration that Chinese hardware can handle serious large-scale workloads. The timing matters. Export controls from the US have restricted access to NVIDIA’s most advanced chips for Chinese buyers, pushing companies like Z.ai to build around what they have. What the model actually does GLM-5.3-Flash uses a hybrid architecture combining sparse and linear attention, which is a design choice that cuts compute requirements significantly. The model carries 320 billion total parameters but activates only 18 billion per token, meaning it draws on a massive knowledge base without paying the full computational cost on every inference. The context window sits at 1 million tokens, placing it among the longest available in any open model. Native support for image and video inputs makes it genuinely multimodal from the architecture level up, rather than as a bolted-on capability. On the DeepSWE v1.1 coding benc...

Read Entire Article