Wan3.0-Video
Wan 3.0 is Alibaba's unified video generation model, and unification is the point: one endpoint replaces the separate text-to-video, image-to-video, reference-to-video, and video-editing models of the previous generation. It produces clips from two to thirty seconds at up to 1080p, with audio generated in the same pass rather than added afterwards. Its distinguishing capability is the breadth of what it will accept as a reference — not only text, images, audio, and video, but documents, spreadsheets, slide decks, and web pages. Hand it a product spec or a quarterly deck and the output is a finished clip rather than a storyboard.

Wan 3.0 Video
Alibaba Tongyi Lab's unified video generation model. One model covering text-to-video, image-to-video, reference-driven generation, and video editing — capabilities that required four separate models in the previous generation.
What Makes It Different
Thirty seconds in a single pass. Not stitched, not extended — generated as one continuous clip. Roughly double the previous generation's ceiling, and long enough to carry a narrative rather than a single beat.
Audio generated with the picture. Sound is produced in the same pass and on by default, not added as a separate step against a finished clip.
Omni-Reference. The model accepts text, images, audio, and video as references — and, for the
first time in this family, structured documents and web pages. Supported formats include doc,
xls, ppt, pdf, txt, key, pages, numbers, and md, alongside any webpage URL.
That last capability is the one worth thinking about. Alibaba's framing is turning static, text-heavy material into video content: hand it a quarterly deck or a product specification and the output is a finished clip, not a storyboard to work from.
Smart duration. The model can read the prompt and choose an appropriate length, so a simple product shot is not padded to thirty seconds and an ambitious narrative is not cut short.
Specifications
| Duration | 2 to 30 seconds |
| Resolution | Up to 1080p — tiers at 480p, 720p, 1080p |
| Audio | Generated alongside video, on by default |
| Reference inputs | Text, image, audio, video, documents, web pages |
| Document limit | One file or link per request · up to 100 MB · up to 50 pages |
| Developer | Alibaba Tongyi Lab |
Not 4K. Published comparisons frequently list this model at 3840×2160. The first-party specification caps output at 1080p.
No architecture is published — no parameter count, no technical report, no model card, and no open weights.
Two Input Modes, Mutually Exclusive
When using references beyond a plain prompt, the model works in one of two modes. Mixing them fails the request.
All-in-one reference
Up to 10 images, 5 video clips, and 5 audio clips as combined references for appearance, motion, and voice.
References are addressed from within the prompt by position — "Image 1", "Video 2" — numbered separately within each type. This is unusual and worth stating clearly: the linkage between a reference and its role in the scene is expressed in natural language, not in the request structure.
First and last frame
A starting frame and an optional ending frame. The model generates the motion between them.
Use this when you know where a shot begins and ends. Use all-in-one reference when you are describing a scene built from elements.
Capabilities
| Capability | Value |
|---|---|
input_types | text |
output_types | video |
audio_output | Yes — generated with the video |
max_duration_seconds | 30 |
max_resolution | 1080p |
endpoint | /v1/video/generations |
streaming | Not applicable |
requires_prompt | Yes |
The model itself accepts image, audio, video, and document references. Whether those reach it through a given API path depends on that path's request schema — verify against the endpoint documentation before building around them.
Using Wan 3.0 on DEVUP AI
Endpoint: POST https://api.devupai.com/v1/video/generations
This is not the chat completions endpoint. Video generation has its own path, its own request shape, and its own response shape.
Basic generation — Python
import os
import requests
DEVUP_API_KEY = os.environ["DEVUP_API_KEY"]
response = requests.post(
"https://api.devupai.com/v1/video/generations",
headers={
"Authorization": f"Bearer {DEVUP_API_KEY}",
"Content-Type": "application/json",
},
json={
"model": "Qwen/Wan3.0-Video",
"prompt": (
"A slow aerial pan over the Casbah of Algiers at golden hour. "
"White terraced rooftops descend toward the bay. Warm low sunlight, "
"long shadows, gentle haze over the water."
),
},
timeout=900, # generation is synchronous and can run for several minutes
)
response.raise_for_status()
result = response.json()
# Download immediately — the signed URL expires 300 seconds after the response.
video_url = result["data"][0]["url"]
with open("output.mp4", "wb") as handle:
handle.write(requests.get(video_url, timeout=120).content)
print(f"Cost: {result['_devup']['cost_dzd']} DZD")
print(f"Balance: {result['_devup']['balance_dzd']} DZD")The timeout is not optional. Generation is synchronous, and a single request can hold the connection open for several minutes. A default client timeout will abort a request the server is still working on — and you will have paid for the generation regardless.
Avoiding the expiring URL — Python
The signed URL expires 300 seconds after the response is produced, after which the proxy returns HTTP 403. If your pipeline cannot download within that window, request the bytes inline instead.
response = requests.post(
"https://api.devupai.com/v1/video/generations",
headers={
"Authorization": f"Bearer {DEVUP_API_KEY}",
"Content-Type": "application/json",
},
json={
"model": "Qwen/Wan3.0-Video",
"prompt": prompt,
"response_format": "b64_json",
},
timeout=900,
)
response.raise_for_status()
result = response.json()
entry = result["data"][0]
if "b64_json" in entry:
import base64
with open("output.mp4", "wb") as handle:
handle.write(base64.b64decode(entry["b64_json"]))
else:
# Videos above the 20 MB inline limit fall back to a signed URL automatically.
# This is a delivery fallback, not an error — the request is billed normally.
note = result["_devup"].get("note")
print(f"Fell back to URL delivery: {note}")
with open("output.mp4", "wb") as handle:
handle.write(requests.get(entry["url"], timeout=120).content)Handle both shapes. A thirty-second 1080p clip will frequently exceed the 20 MB inline limit, so
code that only reads b64_json will fail on exactly the requests you care most about. Only "url"
and "b64_json" are accepted; any other value returns HTTP 400.
Node.js
import { writeFile } from "node:fs/promises";
const response = await fetch("https://api.devupai.com/v1/video/generations", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.DEVUP_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "Qwen/Wan3.0-Video",
prompt: "A slow aerial pan over a coastal city at golden hour.",
}),
signal: AbortSignal.timeout(900_000), // synchronous generation, minutes not seconds
});
if (!response.ok) {
throw new Error(`Request failed with status ${response.status}`);
}
const result = await response.json();
const videoResponse = await fetch(result.data[0].url);
await writeFile("output.mp4", Buffer.from(await videoResponse.arrayBuffer()));
console.log(`Cost: ${result._devup.cost_dzd} DZD`);cURL
curl -X POST "https://api.devupai.com/v1/video/generations" \
-H "Authorization: Bearer $DEVUP_API_KEY" \
-H "Content-Type: application/json" \
--max-time 900 \
-d '{
"model": "Qwen/Wan3.0-Video",
"prompt": "A slow aerial pan over a coastal city at golden hour, warm low sunlight."
}'Writing Prompts That Work
Video prompts reward different things than text prompts.
Describe motion, not just the scene. A still description produces a still-feeling clip. Say what moves, and how.
Name the camera. "Slow pan", "aerial shot", "handheld close-up", "static wide". The model has no default cinematography; if you do not specify, you get whatever it picks.
Describe the light. Time of day, direction, hardness. This does more for perceived quality than almost any other single instruction.
Keep it focused. An overly complex prompt produces inconsistent results — the model divides its attention rather than executing everything well. One scene, one intention.
Let the model choose the duration when you do not have a fixed requirement. Padding a simple shot to thirty seconds produces filler.
Cost and Billing
The _devup object on every response carries cost_dzd for the request and balance_dzd for what
remains. Read it.
Video generation is billed by duration and resolution, which means the cost of a request is a function of choices made in the request itself. A thirty-second 1080p clip and a two-second 480p clip are not in the same order of magnitude.
Two practical consequences:
Prototype at low resolution. Iterate on the prompt at the smallest tier until the shot is right, then generate the final at the resolution you need. The composition, motion, and timing carry across; only the pixels change.
Do not put an unbounded video endpoint behind a public form. There is no per-request spending limit on this platform. A generation queue with no ceiling and a synchronous multi-minute request is a combination that consumes a balance quickly.
Limitations
- 1080p maximum. Widely repeated 4K claims are not supported by the first-party specification.
- Thirty seconds maximum per generation. Longer sequences require extension across segments.
- Synchronous generation. A request holds the connection for minutes; plan timeouts, retries, and user-facing progress accordingly.
- Signed URLs expire after 300 seconds. Download immediately or request inline bytes.
- Inline delivery caps at 20 MB, above which the response falls back to a URL. Handle both.
- No open weights, no model card, no technical report, and no published parameter count or architecture.
- Not independently benchmarked. Published capability claims are first-party.
- Reference and document inputs depend on the request schema of the path you call. The model supports them; verify your endpoint does before designing around them.
- No per-request spending limit exists on this platform. Bound generation volume in your own application layer.
- Not a safety layer. Apply your own moderation to prompts and to output before publishing generated video.