ModelsQwenQwen3.8-2.4T-A95B
providerQwen /

Qwen3.8-2.4T-A95B

700 DZD in 2100 DZD out 70 DZD cached/ 1M tokens

Qwen3.8-2.4T-A95B is the first Qwen-Max-class model released with open weights — 2.4 trillion parameters across 92 layers and 512 experts, activating 95 billion per token. It is also the most constrained model in this catalogue: text only, thinking mandatory, no fast path. Every response begins with a reasoning block, and there is no parameter that changes that. Reasoning depth is adjustable across three levels, and reasoning context is preserved across turns by default. The weights run to 213 files and roughly 4.9 terabytes, under a custom licence rather than a permissive one.

Publicfp8JSONStreaming
Qwen3.8-2.4T-A95B
Capabilities
ToolsReasoningStructured output
ArchitectureTransformer
Context Window262K

Qwen3.8-2.4T-A95B

The first Qwen-Max-class model released with open weights. Also the most constrained model in this catalogue — read the three restrictions before anything else.


Three Hard Constraints

The model card states all three in one sentence, and none of them is configurable.

Text only. Multimodal inputs are not supported. No images, no video, no audio.

Thinking is mandatory. It requires thinking mode for all interactions, and thinking cannot be disabled.

Every response opens with reasoning. Each one automatically begins with a <think> block before the final output — not as a default you can override, but as the model's operating mode.

What follows practically.

There is no fast path. Classification, routing, extraction, formatting — every one of them pays for a reasoning pass. If your pipeline has a mechanical route, it needs a different model.

max_tokens must cover reasoning plus answer, always. A budget sized for the answer alone produces a reasoning block and nothing else.

Any code that reads the first content block by position will find reasoning. Select by field, not by index.


The Open Weights Are Not the Hosted Model

Documented on the card, and worth understanding before you plan around a capability.

A hosted version exists, built on this same model, with four things these weights do not have:

Open weightsHosted version
Vision input❌✅
Non-thinking mode❌✅
Default context262,1441,000,000
Built-in tools❌✅

The difference has been publicly disputed. The repository's licence file and its Hugging Face metadata tag both carry the hosted version's name, which prompted objections in the model's own discussion threads that features named in the licence are not present in the weights it governs.

Read that as a fact about what you are getting rather than as a judgment. These weights are text-only, thinking-only, and 262,144 tokens. A capability described in coverage of the hosted version may not exist in the model you are calling.


Architecture

Total parameters2.4T
Activated per token95B
Layers92
Experts512 routed, top-10
AttentionHybrid — full attention plus Gated DeltaNet
Shared expertsPresent
Speculative decodingOptional multi-token prediction
Context262,144 tokens
FoundationArchitectural basis of the previous generation

Ten experts of 512 — under two percent of the routed pool per token.

Ninety-five billion activated is nonetheless the largest activation count in this catalogue. Sparsity here is not about making the model small; it is about making 2.4 trillion parameters runnable at all.

The decoder combines four things: full attention, Gated DeltaNet linear attention, routed experts, and shared experts — with optional multi-token prediction on top.

One implementation detail worth noting: Gated DeltaNet state tensors require FP32 precision and keep that in the checkpoint contract. A quantisation that treats every tensor uniformly will break them.


Four and a Half Terabytes

The number that decides whether self-hosting is a conversation you can have.

213 weight files. Approximately 4,892 GB.

That is the full-precision release. For comparison within the same model's own ecosystem: an aggressive one-bit community quantisation brings the package to roughly 397 GB — about one twelfth — and it requires a purpose-built loader that streams routed expert tensors from disk on demand rather than holding them in memory.

In that package, layers 0 through 91 carry their routed tensors at roughly one bit, with 512 experts packed into each tensor, and only layer 92 kept at a higher precision.

Through an API none of this concerns you. It is here because it explains the model's position: this is not a model most organisations will run themselves, and calling it is the realistic path.


Reasoning Effort

Three levels, with xhigh as the default.

LevelFor
xhigh (default)Complex tasks demanding thorough analysis
mediumBalanced work
lowShorter deliberation

The default is the deepest setting, and on a model where reasoning cannot be turned off at all, that compounds.

Set it explicitly on every path. xhigh on a task that does not need it produces a long reasoning block, a long wait, and the same answer.

And preserve_thinking is enabled by default for all interactions — reasoning context from previous turns is retained rather than discarded.

Those two settings multiply across a session. Preserved reasoning at xhigh accumulates into the input on every subsequent turn. Lowering the effort shortens not only the current turn but everything the conversation carries forward.


Sampling Parameters

Qwen publish a complete recommended set, which is unusual — most cards give two values.

ParameterValue
temperature1.0
top_p0.95
top_k20
min_p0.0
presence_penalty0.0
repetition_penalty1.0

The card notes that support varies by inference framework. Several of these — top_k, min_p, presence_penalty — are not universally honoured through an OpenAI-compatible interface.

The two zeros are worth reading as deliberate. A presence penalty and a repetition penalty both set to neutral, on a model that produces very long reasoning traces, says Qwen found the model does not need them — and that adding them would interfere with reasoning that legitimately revisits the same idea.


Specifications

Model IDQwen/Qwen3.8-2.4T-A95B
Total parameters2.4T
Activated95B
Layers92
Experts512 routed, top-10
Context window262,144 tokens
InputText only
OutputText
ThinkingMandatory — cannot be disabled
Reasoning effortlow, medium, xhigh (default)
Preserve thinkingEnabled by default
Weight files213, ~4,892 GB
LicenceCustom — not Apache or MIT
DeveloperQwen Team, Alibaba

The licence is custom. Most models in this family ship Apache 2.0; this one does not. Read it against your deployment rather than assuming the family's usual terms carry over.

An FP8 checkpoint is published by Qwen, alongside community NVFP4, GGUF, and sub-one-bit builds.


Capabilities

CapabilityValue
input_typestext
output_typestext
image_inputNot supported
video_inputNot supported
context_window262144
reasoningMandatory
reasoning_effortlow, medium, xhigh
thinking_disableNot supported
preserve_thinkingEnabled by default
reasoning_fieldreasoning_content — separate from content
streamingSupported
tool_callingSupported
structured_outputSupported
requires_promptYes — text prompt required

Using Qwen3.8-2.4T-A95B on DEVUP AI

Base URL: https://api.devupai.com/v1 · Model ID: Qwen/Qwen3.8-2.4T-A95B

Python

PYTHON
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEVUP_API_KEY"],
    base_url="https://api.devupai.com/v1",
)

response = client.chat.completions.create(
    model="Qwen/Qwen3.8-2.4T-A95B",
    messages=[
        {"role": "user", "content": "Hello world!"}
    ],
    max_tokens=1024,
)

print(response.choices[0].message.content)

Node.js

JAVASCRIPT
import DevupAI from "devupai";

const client = new DevupAI({
  apiKey: process.env.DEVUP_API_KEY,
});

async function main() {
  const response = await client.chat.completions.create({
    model: "Qwen/Qwen3.8-2.4T-A95B",
    messages: [{ role: "user", content: "Hello world!" }],
    max_tokens: 1024,
  });

  console.log(response.choices[0].message.content);
}

main();

cURL

BASH
curl -X POST "https://api.devupai.com/v1/chat/completions" \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3.8-2.4T-A95B",
    "messages": [
      { "role": "user", "content": "Hello world!" }
    ],
    "max_tokens": 1024
  }'

At Qwen's Published Settings

Their full recommended parameter set, with the effort level made explicit.

PYTHON
response = client.chat.completions.create(
    model="Qwen/Qwen3.8-2.4T-A95B",
    messages=[
        {"role": "user", "content": hard_question}
    ],
    temperature=1.0,
    top_p=0.95,
    max_tokens=32768,
    extra_body={
        "top_k": 20,
        "min_p": 0.0,
        "chat_template_kwargs": {"reasoning_effort": "medium"},
    },
)

choice = response.choices[0]
message = choice.message

trace = getattr(message, "reasoning_content", None)

if choice.finish_reason == "length":
    raise ValueError(
        f"hit the ceiling at {response.usage.completion_tokens:,} tokens "
        f"with {len(trace or ''):,} characters of reasoning"
    )

print(message.content)

effort is set explicitly rather than inherited. The default is the deepest level, and on a model that always reasons, that is the most expensive configuration available.

The finish_reason check earns its place here more than anywhere. Reasoning is mandatory and shares the budget — a truncated response can contain a complete chain of thought and no answer, which reads as an empty result rather than as an error.

32,768 output tokens is a starting point, not a generous one, on a model whose reasoning cannot be switched off.


Budgeting for Mandatory Reasoning

The measurement worth taking before you set a production ceiling.

PYTHON
def measure(prompt: str, effort: str) -> tuple[int, int]:
    """Return reasoning characters and total completion tokens at a given effort."""
    response = client.chat.completions.create(
        model="Qwen/Qwen3.8-2.4T-A95B",
        messages=[{"role": "user", "content": prompt}],
        temperature=1.0,
        top_p=0.95,
        max_tokens=65536,
        extra_body={"chat_template_kwargs": {"reasoning_effort": effort}},
    )

    message = response.choices[0].message
    trace = getattr(message, "reasoning_content", "") or ""

    return len(trace), response.usage.completion_tokens


for effort in ("low", "medium", "xhigh"):
    chars, tokens = measure(your_prompt, effort)
    print(f"{effort:>6}  {chars:>8,} reasoning chars  {tokens:>7,} total tokens")

Run this across twenty real prompts before choosing a default. The ratio between reasoning and answer is the number that decides your output budget, and it is workload-specific rather than universal.

Watch what low costs you in quality, not just what it saves. On a model where reasoning is the mechanism rather than an option, the lowest setting is still a reasoning pass — it is shorter, not absent.


Multi-Turn Sessions

Preserved thinking is on by default, which changes how a conversation grows.

PYTHON
conversation = []


def turn(text: str, *, effort: str = "medium") -> str:
    """Send one turn and report how the input is growing."""
    conversation.append({"role": "user", "content": text})

    response = client.chat.completions.create(
        model="Qwen/Qwen3.8-2.4T-A95B",
        messages=conversation,
        temperature=1.0,
        top_p=0.95,
        max_tokens=32768,
        extra_body={"chat_template_kwargs": {"reasoning_effort": effort}},
    )

    message = response.choices[0].message
    conversation.append(message)  # append as returned — reasoning travels with it

    usage = response.usage
    print(f"in {usage.prompt_tokens:>8,} · out {usage.completion_tokens:>7,}")

    return message.content

Append the message object as returned, not a rebuilt one. Reasoning travels with it, and reconstructing from content discards what preservation exists to keep.

Watch the input column climb. It rises faster than the visible conversation does — that gap is the preserved reasoning, billed as input on every turn, and at xhigh it climbs considerably faster than at low.

262,144 tokens fills sooner than a turn count suggests when every turn contributes a reasoning block to the history.


Self-Hosting

Realistic only at serious scale, and three details matter if you attempt it.

213 weight files, roughly 4.9 TB in the full-precision release.

Gated DeltaNet state tensors require FP32. A uniform quantisation pass across all tensors breaks the checkpoint contract — the published recipes handle this, a naïve one will not.

Community quantisations reach extreme compression, down to roughly 397 GB at approximately one bit for the routed experts, with only the final layer kept higher. Those packages are not conventional GGUF and need a purpose-built loader that streams expert tensors from disk per layer; a stock build cannot read them.

An official FP8 checkpoint exists from Qwen, and community NVFP4 builds follow an experts-only recipe.

Speculative decoding drafts have been published separately by third parties, trained against specific quantisations of this model.


Where It Fits

The hardest reasoning problems, where mandatory deliberation is what you wanted and 95 billion activated parameters is what the problem needs.

Long-horizon agentic work, which Qwen name as a target — carrying complex multi-step tasks through to completion.

Coding, professional work, and research, the three areas named for comprehensive improvement.

Long-context analysis at 262,144 tokens.

Not for anything mechanical. Classification, routing, extraction, and formatting all pay for a reasoning pass they do not need. Those belong on a smaller model in this family.

Not for vision or video. Text only — the multimodal member of this generation is the 27B dense model.

Not for latency-sensitive paths. There is no fast mode, and the default is the deepest one.

Not for most self-hosting. Four and a half terabytes is a data-centre decision.


Practical Notes

Set reasoning_effort explicitly on every request. The default is the deepest level and thinking cannot be switched off.

Size max_tokens for reasoning plus answer, always — there is no configuration where reasoning is absent.

Check finish_reason on every call.

Use Qwen's full published sampling set, and expect some parameters to be unsupported depending on the serving stack.

Append assistant messages as returned; preserved thinking travels with the object.

Track input growth across a session — preserved reasoning compounds with the effort setting.

Route mechanical work to a smaller model in this family.

Read the licence. It is custom, not the permissive one used elsewhere in this family.


Limitations

Text only. No image, audio, or video input. The multimodal member of this generation is a separate, much smaller model.

Thinking cannot be disabled. Every request reasons, at minimum cost that no parameter removes.

The default effort is the deepest level, which on a mandatory-reasoning model is the most expensive configuration available.

Preserved thinking grows input tokens on every turn, and compounds with the effort setting.

262,144 context — a quarter of what the hosted version of this model provides by default.

Four features present in the hosted version are absent here: vision, non-thinking mode, the million-token window, and built-in tools.

A custom licence, not Apache or MIT. Read it before commercial deployment.

Self-hosting is a data-centre proposition. 213 files, ~4.9 TB at full precision.

Sampling parameter support varies by framework, and Qwen say so.

Reasoning traces are working notes. Unpolished, sometimes exploring abandoned branches, and — at the default effort on a model that cannot stop reasoning — prone to length that exceeds the value it adds.