Modelszai-orgGLM-5.3-Flash
providerzai-org /

GLM-5.3-Flash

60 DZD in 200 DZD out 12 DZD cached/ 1M tokens

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, and it was built from a new base model rather than adapted from an existing one. It carries 320B total parameters but activates only 18B per token, routing each token through 8 of 288 experts across 45 layers. Its defining architectural choice is a hybrid attention stack combining linear and sparse attention — a first for the series — which cuts the serving cost of very long inputs while keeping long-context accuracy intact. Trained on a 30-trillion-token multimodal corpus, it accepts images alongside text and holds a full one-million-token context. Released under the MIT license, it is the model to reach for on DEVUP AI when your input is visual, very long, or both.

Publicfp8JSON
GLM-5.3-Flash
Capabilities
ToolsVisionReasoningStructured outputVideo
ArchitectureTransformer
Context Window1M

Use GLM-5.3-Flash in VS Code, Cursor and Antigravity

DEVUP Code, the DEVUP AI extension, is free to install. Requests are paid from your DEVUP balance.

Explore DEVUP Code

GLM-5.3-Flash

Overview

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. It is a Mixture-of-Experts model with 320B total parameters of which 18B activate per token, and it supports a one-million-token context window with image input alongside text.

It is not an adaptation of the previous generation. The base model was trained from scratch, with the architecture and training recipe redesigned around the same goal the parameter split implies: more capability per unit of compute.

The defining architectural choice is a hybrid attention stack combining linear and sparse attention — a first for the GLM series. Linear-attention layers handle the bulk of the sequence at low cost while sparse attention layers preserve precise retrieval, which is what makes a million-token window affordable to serve without degrading long-context accuracy.


At a Glance

FieldValue
Model TypeMultimodal Mixture-of-Experts transformer
Total Parameters320B
Activated Parameters18B per token
Experts8 active of 288
Layers45 (language model)
Context Window1,048,576 tokens (1M)
PrecisionNative FP8
ModalityText + image in → text out
Tool CallingSupported
Pre-training30T-token multimodal corpus
LanguagesEnglish, Chinese
LicenseMIT

Architecture

ComponentDetail
SparsityMoE — 320B total, 18B activated per token, 8 of 288 experts routed
Depth45 layers in the language model
AttentionHybrid — KDA linear-attention layers interleaved with NoPE sparse MLA layers
Residual pathManifold-Constrained Hyper-Connections (mHC)
Speculative decodingOne Multi-Token Prediction (MTP) draft layer in the checkpoint
PrecisionNative FP8 weights

Hybrid attention is the piece that matters operationally. Pure attention scales quadratically and pure linear attention loses precision on retrieval; interleaving the two gives most of the cost reduction of the latter while keeping the exactness of the former where it counts. This is what turns a million-token window from a specification into a usable workflow.

mHC strengthens the conventional residual connection, improving scaling efficiency and the stability of signal propagation across layers.

MTP is a speculative decoding layer shipped inside the checkpoint rather than as a separate draft model. The practical effect is lower generation latency.


Capabilities

CapabilityValue
input_typestext, image
output_typestext
context_window1048576
streamingSupported
tool_callingSupported
reasoning_fieldreasoning_content — separate from content
requires_promptYes — text prompt required, image optional

Recommended Use Cases

  • Document and screenshot understanding — the combination of native image input and a million-token window means whole scanned documents, long report exports, or full UI captures can be reasoned over in one pass.
  • Visual question answering over large collections — many images plus their surrounding text held together in a single context.
  • Coding agents and long-horizon engineering — strong terminal and repository-level agent results, with the context to hold an entire codebase.
  • Whole-corpus analysis — contracts, logs, specifications; the hybrid attention design is specifically what makes the full window practical rather than nominal.
  • High-throughput production traffic — 18B activated parameters plus in-checkpoint speculative decoding is what makes this viable as a default rather than an escalation tier.
  • Mixed text-and-visual pipelines — workflows that would otherwise need a separate vision model and a separate text model wired together.

Using GLM-5.3-Flash on DEVUP AI

Base URL: https://api.devupai.com/v1 · Model ID: zai-org/GLM-5.3-Flash

Quick start — cURL

BASH
curl https://api.devupai.com/v1/chat/completions \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org/GLM-5.3-Flash",
    "messages": [
      {
        "role": "user",
        "content": "Two services write to the same row without a transaction. Walk through the failure modes in order of likelihood and propose the smallest fix for each."
      }
    ],
    "temperature": 1.0,
    "top_p": 0.95,
    "max_tokens": 16384
  }'

Image input — Python

The capability that separates this model from the rest of its size class.

PYTHON
import os
import base64
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEVUP_API_KEY"],
    base_url="https://api.devupai.com/v1",
)

with open("invoice.png", "rb") as handle:
    encoded = base64.b64encode(handle.read()).decode("utf-8")

response = client.chat.completions.create(
    model="zai-org/GLM-5.3-Flash",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/png;base64,{encoded}"},
                },
                {
                    "type": "text",
                    "text": (
                        "Extract every line item as JSON with fields: description, "
                        "quantity, unit_price, line_total. Report the currency exactly as "
                        "printed. If a field is illegible, use null rather than guessing."
                    ),
                },
            ],
        }
    ],
    temperature=1.0,
    top_p=0.95,
    max_tokens=16384,
)

print(response.choices[0].message.content)

Instructing the model to return null for illegible fields rather than inferring them is not optional in document work. A plausible invented figure is worse than a missing one, because nothing downstream can detect it.

Node.js — DEVUP AI SDK

BASH
npm install devupai
JAVASCRIPT
import DevupAI from "devupai";

const client = new DevupAI({
  apiKey: process.env.DEVUP_API_KEY,
});

const response = await client.chat.completions.create({
  model: "zai-org/GLM-5.3-Flash",
  messages: [
    {
      role: "system",
      content:
        "You are a staff engineer reviewing a migration. Report only defects that would " +
        "cause data loss or downtime, each with a severity and the smallest safe fix.",
    },
    { role: "user", content: migrationPlan },
  ],
  temperature: 1.0,
  top_p: 0.95,
  max_tokens: 16384,
});

console.log(response.choices[0].message.content);

Whole-corpus analysis — Python

PYTHON
with open("repository_bundle.txt", encoding="utf-8") as handle:
    corpus = handle.read()

response = client.chat.completions.create(
    model="zai-org/GLM-5.3-Flash",
    messages=[
        {
            "role": "system",
            "content": (
                "You are auditing a codebase. Identify every place a database write occurs "
                "outside a transaction. Cite the file and function for each finding, and "
                "report nothing you cannot cite."
            ),
        },
        {"role": "user", "content": corpus},
    ],
    temperature=1.0,
    top_p=1.0,
    max_tokens=65536,
)

print(response.choices[0].message.content)

Cross-file reasoning of this kind is what the full window buys that a retrieval layer cannot: chunking discards exactly the relationships being searched for.

Reading the reasoning trace

Reasoning is separated from the answer rather than embedded in it. Read it explicitly, and null-check it — not every model on the platform populates the field.

PYTHON
message = response.choices[0].message

reasoning = getattr(message, "reasoning_content", None)
if reasoning:
    # Log it, do not show it. Reasoning traces are intermediate, not conclusions.
    logger.debug("trace length: %d chars", len(reasoning))

print(message.content)

Never concatenate reasoning_content into content before parsing or display. Doing so breaks JSON parsing on structured-output paths and shows users an unpolished draft of an answer they never asked to see.

Streaming with usage

PYTHON
stream = client.chat.completions.create(
    model="zai-org/GLM-5.3-Flash",
    messages=[{"role": "user", "content": "Design a retry policy for a webhook delivery system."}],
    temperature=1.0,
    top_p=0.95,
    max_tokens=16384,
    stream=True,
    stream_options={"include_usage": True},
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print(f"\n\nTokens — in: {chunk.usage.prompt_tokens}, out: {chunk.usage.completion_tokens}")

Image inputs consume a substantial number of input tokens. On visual workloads stream_options.include_usage is the only way to see how many.

Delegating access with a scoped JWT

Issue a token restricted to this model with an expiry and a spending limit instead of sharing your API key with an unattended agent or a public-facing upload endpoint:

BASH
curl -X POST "https://api.devupai.com/v1/scoped-jwt" \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "api_key_name": "auto",
    "models": ["zai-org/GLM-5.3-Flash"],
    "expires_delta": 3600,
    "spending_limit": 200
  }'

The returned token is used exactly like an API key in the Authorization header. Requests for any other model, or past the expiry or spending limit, are rejected — a hard ceiling on what an unbounded job can consume.


Recommended Generation Parameters

Z.ai publishes different sampling settings per evaluated workload rather than one universal pair. The pattern is clear enough to use as guidance.

Workloadtemperaturetop_p
Reasoning with tools1.00.95
Visual tasks1.00.95
Repository-scale code generation1.01.0
Terminal and agent work1.01.0
Software-engineering agents0.951.0

For image inputs, Z.ai resizes so the shorter side is at least roughly 1,500 pixels in its own evaluations. Downscaling below that is the most common cause of avoidable extraction errors on dense documents.


Benchmark Results

Independently listed evaluation results for this checkpoint.

BenchmarkScore
Terminal-Bench 2.184.3
DeepSWE63.4
HLE (with tools)55.3

Z.ai reports that this model outperforms the previous generation in its series across benchmarks and real-world workloads, and publishes a broader comparison table in its own materials.


Best Practices

  • Send images at sufficient resolution. Roughly 1,500 pixels on the shorter side is the reference point; below that, dense text in documents starts to fail silently.
  • Instruct the model to return null rather than guess on any extraction task. An invented value is undetectable downstream.
  • Match sampling to the workload. top_p of 0.95 for reasoning and visual work, 1.0 for code generation and agent loops.
  • Use the context window instead of building retrieval where the corpus fits. The hybrid attention design exists precisely so that this is affordable.
  • Read reasoning_content as a separate field. Do not merge it into content, and null-check it.
  • Watch input token counts on visual requests. Images are not free context, and a batch of high-resolution pages consumes far more than the surrounding text.
  • Bound every agent loop with an iteration ceiling and a scoped token.

Limitations

  • Text output only. The model accepts visual input but does not generate images.
  • Two evaluated languages. English and Chinese are the declared languages; other languages may work but are not covered by the published evaluation.
  • Reasoning traces are not conclusions. Content in reasoning_content may be unpolished or contradict the final answer.
  • Benchmark figures are harness-dependent. Published agent scores depend heavily on the scaffold, timeout, and context budget used to produce them.
  • Not a safety layer. Apply your own moderation and validation before acting on model output in a production system.