How to Use the Wan 3.0 Omni Reference API: Omni Reference integration guide (2026)
Learn how to call the Wan 3.0 Omni Reference API on NamiFusion: parameters, pricing (from ≈$1.00) and runnable copy-paste code — try it in the playground first.
Wan 3.0 Omni Reference is an Omni Reference model by Alibaba: Generates videos up to 30 seconds with native audio from multimodal references — images, video, audio, documents, and web pages. Reproduces characters, props, and spatial relationships from your references at pixel-level fidelity, with a first/last-frame mode for precise control. On NamiFusion you can try it in the playground (exact quote shown before you submit), then call it through one unified REST API — a typical run costs about $1.00.
TL;DR
· Task type: Omni Reference, provided by Alibaba
· Pricing: about $1.00 per typical run ($1 = 100 credits), pay-as-you-go
· Getting started: playground (exact quote before submit) → create an API key → one POST request
· Tunable parameters: 12 (see the table below)
· No platform watermark; commercial use allowed (per the Terms of Service)
What can you build with Wan 3.0 Omni Reference?
Content automation: wire this model into your workflow and call it on demand through one unified API.
How do you call this API?
1. Sign up on NamiFusion and create a key on the API Keys page (signup bonus credits can only be used on selected models).
2. Submit a task with the cURL example below, passing your key as a Bearer token.
3. Poll the task status with the returned task_uuid and read the output URLs when completed.
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/alibaba/wan-3.0/video" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"resolution": "1080P",
"aspect_ratio": "adaptive",
"duration": 5,
"generate_audio": true,
"enable_thinking": false
}
}'
# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"What parameters does it take?
prompt (Prompt): optional
images (Reference Images): optional
audios (Reference Audio): optional
first_frame (First Frame): optional
last_frame (Last Frame): optional
resolution (Resolution): optional, default "1080P"
aspect_ratio (Aspect Ratio): optional, default "adaptive"
duration (Duration (seconds)): optional, default 5
generate_audio (Generate Audio): optional, default true
enable_thinking (Thinking Mode): optional, default false
file (Reference Document): optional
link (Reference Web Page): optional
How do you write prompts for Wan 3.0 Omni Reference?
1. Subject first: lead with the actor and action, then add environment, lighting and style ("an astronaut sprinting through a neon street in the rain, cinematic, shallow depth of field").
2. Concrete nouns beat stacked adjectives: "85mm portrait lens, golden-hour backlight" works far better than "very beautiful, ultra high quality".
3. Change one variable per iteration: pin the seed to reproduce a result, tweak a single phrase at a time so you know what actually moved the output.
4. State negatives positively: instead of "no blur", ask for "sharp focus, crisp details".
How much does it cost?
Pay-per-use in credits. A typical run costs about $1.00 ($1 = 100 credits); the exact price varies with parameters like resolution or duration, and the playground shows the exact quote before you submit.
Wan 3.0 Omni Reference
Turn images, video, audio, documents, and web pages into one video
Alibaba Wan 3.0 Omni Reference accepts five kinds of reference input at once and generates a single video of up to 30 seconds with a native audio track. It reproduces characters, props, and spatial relationships from your references at pixel-level fidelity, supports first/last-frame control for precise framing, and can decide the final length on its own. Ready-to-use REST inference API, no cold starts.
🚀 Key Features
- Five reference modalities: Up to 10 reference images, 5 reference videos, and 5 audio clips, plus one document or one web link — combine them freely.
- Pixel-level identity preservation: Characters, props, wardrobe, and spatial relationships stay consistent across new shots, with no extra training.
- Native audio: Outputs video with sound by default — no separate dubbing pass.
- First/last-frame control: Pin the opening frame, the closing frame, or both.
- Model-decided duration: Hand the length to the model and let the content determine it (up to 30 seconds).
- Thinking mode: Parses reference documents, web pages, and complex imagery before generating.
🛠️ Technical Specifications
| Parameter | Description |
|---|---|
| Model Architecture | Alibaba Wan 3.0 |
| Reference Images | Up to 10 |
| Reference Videos | Up to 5 (their duration is billed) |
| Reference Audio | Up to 5 |
| Reference Document / Web Page | Document ≤100MB and ≤50 pages, or one public URL; mutually exclusive, requires thinking mode |
| Output Format | Video (MP4, audio included by default) |
| Resolution | 480P / 720P / 1080P (default 1080P) |
| Aspect Ratio | adaptive / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 |
| Duration | 2–30 seconds (default 5), or -1 to let the model decide; with video input, input plus output must not exceed 30 seconds |
| Timeout | 300 seconds |
💰 Pricing
| Resolution | Pricing formula |
|---|---|
| 480P | 20 credits × total_duration |
| 720P | 40 credits × total_duration |
| 1080P | 80 credits × total_duration |
total_duration = generated video duration + billed duration of reference video inputs (summed across all reference clips).
How model-decided duration is billed: setting duration to -1 lets the model choose the length, which cannot be known at request time, so it is billed at the 30-second maximum. For example, 1080P with model-decided duration and no reference video costs 2400 credits per call. If you already know roughly how long the clip should be, passing an explicit number is cheaper.
FAQ
Can I try it for free first?+
Signup bonus credits can only be used on selected models, and eligibility may change. The playground shows the exact quote before submission.
Can I use the output commercially? Is there a watermark?+
Yes, outputs are yours to use commercially, and NamiFusion adds no platform watermark — subject to the Terms of Service and applicable law.
What is the content policy?+
NamiFusion is an unfiltered engine: no arbitrary restrictions and NSFW-friendly, but illegal content and non-consensual use of real people are prohibited — see the Terms of Service.
How long does a Wan 3.0 Omni Reference task take?+
It depends on parameters — images usually take seconds, videos tens of seconds to minutes. The API returns a task_uuid for polling.
What are the resolution and duration limits?+
See the parameter table above — each model lists its available resolution and duration options there and in the playground.