Qwen3.8-2.4T-A95B
Qwen3.8-2.4T-A95B is the first Qwen-Max-class model released with open weights — 2.4 trillion parameters across 92 layers and 512 experts, activating 95 billion per token. It is also the most constrained model in this catalogue: text only, thinking mandatory, no fast path. Every response begins with a reasoning block, and there is no parameter that changes that. Reasoning depth is adjustable across three levels, and reasoning context is preserved across turns by default. The weights run to 213 files and roughly 4.9 terabytes, under a custom licence rather than a permissive one.

Qwen3.8-2.4T-A95B
The first Qwen-Max-class model released with open weights. Also the most constrained model in this catalogue — read the three restrictions before anything else.
Three Hard Constraints
The model card states all three in one sentence, and none of them is configurable.
Text only. Multimodal inputs are not supported. No images, no video, no audio.
Thinking is mandatory. It requires thinking mode for all interactions, and thinking cannot be disabled.
Every response opens with reasoning. Each one automatically begins with a <think> block before
the final output — not as a default you can override, but as the model's operating mode.
What follows practically.
There is no fast path. Classification, routing, extraction, formatting — every one of them pays for a reasoning pass. If your pipeline has a mechanical route, it needs a different model.
max_tokens must cover reasoning plus answer, always. A budget sized for the answer alone
produces a reasoning block and nothing else.
Any code that reads the first content block by position will find reasoning. Select by field, not by index.
The Open Weights Are Not the Hosted Model
Documented on the card, and worth understanding before you plan around a capability.
A hosted version exists, built on this same model, with four things these weights do not have:
| Open weights | Hosted version | |
|---|---|---|
| Vision input | ❌ | ✅ |
| Non-thinking mode | ❌ | ✅ |
| Default context | 262,144 | 1,000,000 |
| Built-in tools | ❌ | ✅ |
The difference has been publicly disputed. The repository's licence file and its Hugging Face metadata tag both carry the hosted version's name, which prompted objections in the model's own discussion threads that features named in the licence are not present in the weights it governs.
Read that as a fact about what you are getting rather than as a judgment. These weights are text-only, thinking-only, and 262,144 tokens. A capability described in coverage of the hosted version may not exist in the model you are calling.
Architecture
| Total parameters | 2.4T |
| Activated per token | 95B |
| Layers | 92 |
| Experts | 512 routed, top-10 |
| Attention | Hybrid — full attention plus Gated DeltaNet |
| Shared experts | Present |
| Speculative decoding | Optional multi-token prediction |
| Context | 262,144 tokens |
| Foundation | Architectural basis of the previous generation |
Ten experts of 512 — under two percent of the routed pool per token.
Ninety-five billion activated is nonetheless the largest activation count in this catalogue. Sparsity here is not about making the model small; it is about making 2.4 trillion parameters runnable at all.
The decoder combines four things: full attention, Gated DeltaNet linear attention, routed experts, and shared experts — with optional multi-token prediction on top.
One implementation detail worth noting: Gated DeltaNet state tensors require FP32 precision and keep that in the checkpoint contract. A quantisation that treats every tensor uniformly will break them.
Four and a Half Terabytes
The number that decides whether self-hosting is a conversation you can have.
213 weight files. Approximately 4,892 GB.
That is the full-precision release. For comparison within the same model's own ecosystem: an aggressive one-bit community quantisation brings the package to roughly 397 GB — about one twelfth — and it requires a purpose-built loader that streams routed expert tensors from disk on demand rather than holding them in memory.
In that package, layers 0 through 91 carry their routed tensors at roughly one bit, with 512 experts packed into each tensor, and only layer 92 kept at a higher precision.
Through an API none of this concerns you. It is here because it explains the model's position: this is not a model most organisations will run themselves, and calling it is the realistic path.
Reasoning Effort
Three levels, with xhigh as the default.
| Level | For |
|---|---|
xhigh (default) | Complex tasks demanding thorough analysis |
medium | Balanced work |
low | Shorter deliberation |
The default is the deepest setting, and on a model where reasoning cannot be turned off at all, that compounds.
Set it explicitly on every path. xhigh on a task that does not need it produces a long
reasoning block, a long wait, and the same answer.
And preserve_thinking is enabled by default for all interactions — reasoning context from
previous turns is retained rather than discarded.
Those two settings multiply across a session. Preserved reasoning at xhigh accumulates into the
input on every subsequent turn. Lowering the effort shortens not only the current turn but everything
the conversation carries forward.
Sampling Parameters
Qwen publish a complete recommended set, which is unusual — most cards give two values.
| Parameter | Value |
|---|---|
temperature | 1.0 |
top_p | 0.95 |
top_k | 20 |
min_p | 0.0 |
presence_penalty | 0.0 |
repetition_penalty | 1.0 |
The card notes that support varies by inference framework. Several of these — top_k, min_p,
presence_penalty — are not universally honoured through an OpenAI-compatible interface.
The two zeros are worth reading as deliberate. A presence penalty and a repetition penalty both set to neutral, on a model that produces very long reasoning traces, says Qwen found the model does not need them — and that adding them would interfere with reasoning that legitimately revisits the same idea.
Specifications
| Model ID | Qwen/Qwen3.8-2.4T-A95B |
| Total parameters | 2.4T |
| Activated | 95B |
| Layers | 92 |
| Experts | 512 routed, top-10 |
| Context window | 262,144 tokens |
| Input | Text only |
| Output | Text |
| Thinking | Mandatory — cannot be disabled |
| Reasoning effort | low, medium, xhigh (default) |
| Preserve thinking | Enabled by default |
| Weight files | 213, ~4,892 GB |
| Licence | Custom — not Apache or MIT |
| Developer | Qwen Team, Alibaba |
The licence is custom. Most models in this family ship Apache 2.0; this one does not. Read it against your deployment rather than assuming the family's usual terms carry over.
An FP8 checkpoint is published by Qwen, alongside community NVFP4, GGUF, and sub-one-bit builds.
Capabilities
| Capability | Value |
|---|---|
input_types | text |
output_types | text |
image_input | Not supported |
video_input | Not supported |
context_window | 262144 |
reasoning | Mandatory |
reasoning_effort | low, medium, xhigh |
thinking_disable | Not supported |
preserve_thinking | Enabled by default |
reasoning_field | reasoning_content — separate from content |
streaming | Supported |
tool_calling | Supported |
structured_output | Supported |
requires_prompt | Yes — text prompt required |
Using Qwen3.8-2.4T-A95B on DEVUP AI
Base URL: https://api.devupai.com/v1 · Model ID: Qwen/Qwen3.8-2.4T-A95B
Python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEVUP_API_KEY"],
base_url="https://api.devupai.com/v1",
)
response = client.chat.completions.create(
model="Qwen/Qwen3.8-2.4T-A95B",
messages=[
{"role": "user", "content": "Hello world!"}
],
max_tokens=1024,
)
print(response.choices[0].message.content)Node.js
import DevupAI from "devupai";
const client = new DevupAI({
apiKey: process.env.DEVUP_API_KEY,
});
async function main() {
const response = await client.chat.completions.create({
model: "Qwen/Qwen3.8-2.4T-A95B",
messages: [{ role: "user", content: "Hello world!" }],
max_tokens: 1024,
});
console.log(response.choices[0].message.content);
}
main();cURL
curl -X POST "https://api.devupai.com/v1/chat/completions" \
-H "Authorization: Bearer $DEVUP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3.8-2.4T-A95B",
"messages": [
{ "role": "user", "content": "Hello world!" }
],
"max_tokens": 1024
}'At Qwen's Published Settings
Their full recommended parameter set, with the effort level made explicit.
response = client.chat.completions.create(
model="Qwen/Qwen3.8-2.4T-A95B",
messages=[
{"role": "user", "content": hard_question}
],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={
"top_k": 20,
"min_p": 0.0,
"chat_template_kwargs": {"reasoning_effort": "medium"},
},
)
choice = response.choices[0]
message = choice.message
trace = getattr(message, "reasoning_content", None)
if choice.finish_reason == "length":
raise ValueError(
f"hit the ceiling at {response.usage.completion_tokens:,} tokens "
f"with {len(trace or ''):,} characters of reasoning"
)
print(message.content)effort is set explicitly rather than inherited. The default is the deepest level, and on a model
that always reasons, that is the most expensive configuration available.
The finish_reason check earns its place here more than anywhere. Reasoning is mandatory and
shares the budget — a truncated response can contain a complete chain of thought and no answer, which
reads as an empty result rather than as an error.
32,768 output tokens is a starting point, not a generous one, on a model whose reasoning cannot be switched off.
Budgeting for Mandatory Reasoning
The measurement worth taking before you set a production ceiling.
def measure(prompt: str, effort: str) -> tuple[int, int]:
"""Return reasoning characters and total completion tokens at a given effort."""
response = client.chat.completions.create(
model="Qwen/Qwen3.8-2.4T-A95B",
messages=[{"role": "user", "content": prompt}],
temperature=1.0,
top_p=0.95,
max_tokens=65536,
extra_body={"chat_template_kwargs": {"reasoning_effort": effort}},
)
message = response.choices[0].message
trace = getattr(message, "reasoning_content", "") or ""
return len(trace), response.usage.completion_tokens
for effort in ("low", "medium", "xhigh"):
chars, tokens = measure(your_prompt, effort)
print(f"{effort:>6} {chars:>8,} reasoning chars {tokens:>7,} total tokens")Run this across twenty real prompts before choosing a default. The ratio between reasoning and answer is the number that decides your output budget, and it is workload-specific rather than universal.
Watch what low costs you in quality, not just what it saves. On a model where reasoning is the
mechanism rather than an option, the lowest setting is still a reasoning pass — it is shorter, not
absent.
Multi-Turn Sessions
Preserved thinking is on by default, which changes how a conversation grows.
conversation = []
def turn(text: str, *, effort: str = "medium") -> str:
"""Send one turn and report how the input is growing."""
conversation.append({"role": "user", "content": text})
response = client.chat.completions.create(
model="Qwen/Qwen3.8-2.4T-A95B",
messages=conversation,
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": effort}},
)
message = response.choices[0].message
conversation.append(message) # append as returned — reasoning travels with it
usage = response.usage
print(f"in {usage.prompt_tokens:>8,} · out {usage.completion_tokens:>7,}")
return message.contentAppend the message object as returned, not a rebuilt one. Reasoning travels with it, and
reconstructing from content discards what preservation exists to keep.
Watch the input column climb. It rises faster than the visible conversation does — that gap is
the preserved reasoning, billed as input on every turn, and at xhigh it climbs considerably faster
than at low.
262,144 tokens fills sooner than a turn count suggests when every turn contributes a reasoning block to the history.
Self-Hosting
Realistic only at serious scale, and three details matter if you attempt it.
213 weight files, roughly 4.9 TB in the full-precision release.
Gated DeltaNet state tensors require FP32. A uniform quantisation pass across all tensors breaks the checkpoint contract — the published recipes handle this, a naïve one will not.
Community quantisations reach extreme compression, down to roughly 397 GB at approximately one bit for the routed experts, with only the final layer kept higher. Those packages are not conventional GGUF and need a purpose-built loader that streams expert tensors from disk per layer; a stock build cannot read them.
An official FP8 checkpoint exists from Qwen, and community NVFP4 builds follow an experts-only recipe.
Speculative decoding drafts have been published separately by third parties, trained against specific quantisations of this model.
Where It Fits
The hardest reasoning problems, where mandatory deliberation is what you wanted and 95 billion activated parameters is what the problem needs.
Long-horizon agentic work, which Qwen name as a target — carrying complex multi-step tasks through to completion.
Coding, professional work, and research, the three areas named for comprehensive improvement.
Long-context analysis at 262,144 tokens.
Not for anything mechanical. Classification, routing, extraction, and formatting all pay for a reasoning pass they do not need. Those belong on a smaller model in this family.
Not for vision or video. Text only — the multimodal member of this generation is the 27B dense model.
Not for latency-sensitive paths. There is no fast mode, and the default is the deepest one.
Not for most self-hosting. Four and a half terabytes is a data-centre decision.
Practical Notes
Set reasoning_effort explicitly on every request. The default is the deepest level and thinking
cannot be switched off.
Size max_tokens for reasoning plus answer, always — there is no configuration where reasoning is
absent.
Check finish_reason on every call.
Use Qwen's full published sampling set, and expect some parameters to be unsupported depending on the serving stack.
Append assistant messages as returned; preserved thinking travels with the object.
Track input growth across a session — preserved reasoning compounds with the effort setting.
Route mechanical work to a smaller model in this family.
Read the licence. It is custom, not the permissive one used elsewhere in this family.
Limitations
Text only. No image, audio, or video input. The multimodal member of this generation is a separate, much smaller model.
Thinking cannot be disabled. Every request reasons, at minimum cost that no parameter removes.
The default effort is the deepest level, which on a mandatory-reasoning model is the most expensive configuration available.
Preserved thinking grows input tokens on every turn, and compounds with the effort setting.
262,144 context — a quarter of what the hosted version of this model provides by default.
Four features present in the hosted version are absent here: vision, non-thinking mode, the million-token window, and built-in tools.
A custom licence, not Apache or MIT. Read it before commercial deployment.
Self-hosting is a data-centre proposition. 213 files, ~4.9 TB at full precision.
Sampling parameter support varies by framework, and Qwen say so.
Reasoning traces are working notes. Unpolished, sometimes exploring abandoned branches, and — at the default effort on a model that cannot stop reasoning — prone to length that exceeds the value it adds.