GLM-5.3-Flash
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, and it was built from a new base model rather than adapted from an existing one. It carries 320B total parameters but activates only 18B per token, routing each token through 8 of 288 experts across 45 layers. Its defining architectural choice is a hybrid attention stack combining linear and sparse attention — a first for the series — which cuts the serving cost of very long inputs while keeping long-context accuracy intact. Trained on a 30-trillion-token multimodal corpus, it accepts images alongside text and holds a full one-million-token context. Released under the MIT license, it is the model to reach for on DEVUP AI when your input is visual, very long, or both.

Use GLM-5.3-Flash in VS Code, Cursor and Antigravity
DEVUP Code, the DEVUP AI extension, is free to install. Requests are paid from your DEVUP balance.
GLM-5.3-Flash
Overview
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. It is a Mixture-of-Experts model with 320B total parameters of which 18B activate per token, and it supports a one-million-token context window with image input alongside text.
It is not an adaptation of the previous generation. The base model was trained from scratch, with the architecture and training recipe redesigned around the same goal the parameter split implies: more capability per unit of compute.
The defining architectural choice is a hybrid attention stack combining linear and sparse attention — a first for the GLM series. Linear-attention layers handle the bulk of the sequence at low cost while sparse attention layers preserve precise retrieval, which is what makes a million-token window affordable to serve without degrading long-context accuracy.
At a Glance
| Field | Value |
|---|---|
| Model Type | Multimodal Mixture-of-Experts transformer |
| Total Parameters | 320B |
| Activated Parameters | 18B per token |
| Experts | 8 active of 288 |
| Layers | 45 (language model) |
| Context Window | 1,048,576 tokens (1M) |
| Precision | Native FP8 |
| Modality | Text + image in → text out |
| Tool Calling | Supported |
| Pre-training | 30T-token multimodal corpus |
| Languages | English, Chinese |
| License | MIT |
Architecture
| Component | Detail |
|---|---|
| Sparsity | MoE — 320B total, 18B activated per token, 8 of 288 experts routed |
| Depth | 45 layers in the language model |
| Attention | Hybrid — KDA linear-attention layers interleaved with NoPE sparse MLA layers |
| Residual path | Manifold-Constrained Hyper-Connections (mHC) |
| Speculative decoding | One Multi-Token Prediction (MTP) draft layer in the checkpoint |
| Precision | Native FP8 weights |
Hybrid attention is the piece that matters operationally. Pure attention scales quadratically and pure linear attention loses precision on retrieval; interleaving the two gives most of the cost reduction of the latter while keeping the exactness of the former where it counts. This is what turns a million-token window from a specification into a usable workflow.
mHC strengthens the conventional residual connection, improving scaling efficiency and the stability of signal propagation across layers.
MTP is a speculative decoding layer shipped inside the checkpoint rather than as a separate draft model. The practical effect is lower generation latency.
Capabilities
| Capability | Value |
|---|---|
input_types | text, image |
output_types | text |
context_window | 1048576 |
streaming | Supported |
tool_calling | Supported |
reasoning_field | reasoning_content — separate from content |
requires_prompt | Yes — text prompt required, image optional |
Recommended Use Cases
- Document and screenshot understanding — the combination of native image input and a million-token window means whole scanned documents, long report exports, or full UI captures can be reasoned over in one pass.
- Visual question answering over large collections — many images plus their surrounding text held together in a single context.
- Coding agents and long-horizon engineering — strong terminal and repository-level agent results, with the context to hold an entire codebase.
- Whole-corpus analysis — contracts, logs, specifications; the hybrid attention design is specifically what makes the full window practical rather than nominal.
- High-throughput production traffic — 18B activated parameters plus in-checkpoint speculative decoding is what makes this viable as a default rather than an escalation tier.
- Mixed text-and-visual pipelines — workflows that would otherwise need a separate vision model and a separate text model wired together.
Using GLM-5.3-Flash on DEVUP AI
Base URL: https://api.devupai.com/v1 · Model ID: zai-org/GLM-5.3-Flash
Quick start — cURL
curl https://api.devupai.com/v1/chat/completions \
-H "Authorization: Bearer $DEVUP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "zai-org/GLM-5.3-Flash",
"messages": [
{
"role": "user",
"content": "Two services write to the same row without a transaction. Walk through the failure modes in order of likelihood and propose the smallest fix for each."
}
],
"temperature": 1.0,
"top_p": 0.95,
"max_tokens": 16384
}'Image input — Python
The capability that separates this model from the rest of its size class.
import os
import base64
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEVUP_API_KEY"],
base_url="https://api.devupai.com/v1",
)
with open("invoice.png", "rb") as handle:
encoded = base64.b64encode(handle.read()).decode("utf-8")
response = client.chat.completions.create(
model="zai-org/GLM-5.3-Flash",
messages=[
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{encoded}"},
},
{
"type": "text",
"text": (
"Extract every line item as JSON with fields: description, "
"quantity, unit_price, line_total. Report the currency exactly as "
"printed. If a field is illegible, use null rather than guessing."
),
},
],
}
],
temperature=1.0,
top_p=0.95,
max_tokens=16384,
)
print(response.choices[0].message.content)Instructing the model to return null for illegible fields rather than inferring them is not
optional in document work. A plausible invented figure is worse than a missing one, because
nothing downstream can detect it.
Node.js — DEVUP AI SDK
npm install devupaiimport DevupAI from "devupai";
const client = new DevupAI({
apiKey: process.env.DEVUP_API_KEY,
});
const response = await client.chat.completions.create({
model: "zai-org/GLM-5.3-Flash",
messages: [
{
role: "system",
content:
"You are a staff engineer reviewing a migration. Report only defects that would " +
"cause data loss or downtime, each with a severity and the smallest safe fix.",
},
{ role: "user", content: migrationPlan },
],
temperature: 1.0,
top_p: 0.95,
max_tokens: 16384,
});
console.log(response.choices[0].message.content);Whole-corpus analysis — Python
with open("repository_bundle.txt", encoding="utf-8") as handle:
corpus = handle.read()
response = client.chat.completions.create(
model="zai-org/GLM-5.3-Flash",
messages=[
{
"role": "system",
"content": (
"You are auditing a codebase. Identify every place a database write occurs "
"outside a transaction. Cite the file and function for each finding, and "
"report nothing you cannot cite."
),
},
{"role": "user", "content": corpus},
],
temperature=1.0,
top_p=1.0,
max_tokens=65536,
)
print(response.choices[0].message.content)Cross-file reasoning of this kind is what the full window buys that a retrieval layer cannot: chunking discards exactly the relationships being searched for.
Reading the reasoning trace
Reasoning is separated from the answer rather than embedded in it. Read it explicitly, and null-check it — not every model on the platform populates the field.
message = response.choices[0].message
reasoning = getattr(message, "reasoning_content", None)
if reasoning:
# Log it, do not show it. Reasoning traces are intermediate, not conclusions.
logger.debug("trace length: %d chars", len(reasoning))
print(message.content)Never concatenate reasoning_content into content before parsing or display. Doing so breaks
JSON parsing on structured-output paths and shows users an unpolished draft of an answer they
never asked to see.
Streaming with usage
stream = client.chat.completions.create(
model="zai-org/GLM-5.3-Flash",
messages=[{"role": "user", "content": "Design a retry policy for a webhook delivery system."}],
temperature=1.0,
top_p=0.95,
max_tokens=16384,
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print(f"\n\nTokens — in: {chunk.usage.prompt_tokens}, out: {chunk.usage.completion_tokens}")Image inputs consume a substantial number of input tokens. On visual workloads
stream_options.include_usage is the only way to see how many.
Delegating access with a scoped JWT
Issue a token restricted to this model with an expiry and a spending limit instead of sharing your API key with an unattended agent or a public-facing upload endpoint:
curl -X POST "https://api.devupai.com/v1/scoped-jwt" \
-H "Authorization: Bearer $DEVUP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"api_key_name": "auto",
"models": ["zai-org/GLM-5.3-Flash"],
"expires_delta": 3600,
"spending_limit": 200
}'The returned token is used exactly like an API key in the Authorization header. Requests for
any other model, or past the expiry or spending limit, are rejected — a hard ceiling on what an
unbounded job can consume.
Recommended Generation Parameters
Z.ai publishes different sampling settings per evaluated workload rather than one universal pair. The pattern is clear enough to use as guidance.
| Workload | temperature | top_p |
|---|---|---|
| Reasoning with tools | 1.0 | 0.95 |
| Visual tasks | 1.0 | 0.95 |
| Repository-scale code generation | 1.0 | 1.0 |
| Terminal and agent work | 1.0 | 1.0 |
| Software-engineering agents | 0.95 | 1.0 |
For image inputs, Z.ai resizes so the shorter side is at least roughly 1,500 pixels in its own evaluations. Downscaling below that is the most common cause of avoidable extraction errors on dense documents.
Benchmark Results
Independently listed evaluation results for this checkpoint.
| Benchmark | Score |
|---|---|
| Terminal-Bench 2.1 | 84.3 |
| DeepSWE | 63.4 |
| HLE (with tools) | 55.3 |
Z.ai reports that this model outperforms the previous generation in its series across benchmarks and real-world workloads, and publishes a broader comparison table in its own materials.
Best Practices
- Send images at sufficient resolution. Roughly 1,500 pixels on the shorter side is the reference point; below that, dense text in documents starts to fail silently.
- Instruct the model to return
nullrather than guess on any extraction task. An invented value is undetectable downstream. - Match sampling to the workload.
top_pof 0.95 for reasoning and visual work, 1.0 for code generation and agent loops. - Use the context window instead of building retrieval where the corpus fits. The hybrid attention design exists precisely so that this is affordable.
- Read
reasoning_contentas a separate field. Do not merge it intocontent, and null-check it. - Watch input token counts on visual requests. Images are not free context, and a batch of high-resolution pages consumes far more than the surrounding text.
- Bound every agent loop with an iteration ceiling and a scoped token.
Limitations
- Text output only. The model accepts visual input but does not generate images.
- Two evaluated languages. English and Chinese are the declared languages; other languages may work but are not covered by the published evaluation.
- Reasoning traces are not conclusions. Content in
reasoning_contentmay be unpolished or contradict the final answer. - Benchmark figures are harness-dependent. Published agent scores depend heavily on the scaffold, timeout, and context budget used to produce them.
- Not a safety layer. Apply your own moderation and validation before acting on model output in a production system.