Model Library
Browse and deploy state-of-the-art AI models through the DevUp Gateway.
Browse and deploy state-of-the-art AI models through the DevUp Gateway.
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, and it was built from a new base model rather than adapted from an existing one. It carries 320B total parameters but activates only 18B per token, routing each token through 8 of 288 experts across 45 layers. Its defining architectural choice is a hybrid attention stack combining linear and sparse attention — a first for the series — which cuts the serving cost of very long inputs while keeping long-context accuracy intact. Trained on a 30-trillion-token multimodal corpus, it accepts images alongside text and holds a full one-million-token context. Released under the MIT license, it is the model to reach for on DEVUP AI when your input is visual, very long, or both.
