How to Use the Wan 3.0 Prime API: Omni Reference integration guide (2026)

2 min · Updated 2026-07-11 · By the NamiFusion team

Learn how to call the Wan 3.0 Prime API on NamiFusion: parameters, pricing (from ≈$1.25) and runnable copy-paste code — try it in the playground first.

Wan 3.0 Prime is an Omni Reference model by Alibaba: Supports multimodal references including multiple images, video, and audio to generate videos with rich motion and synchronized audio-visual effects. Accurately preserves the appearance, motion characteristics, and visual style of referenced subjects, while supporting flexible combinations of references for versatile content creation and shot control. On NamiFusion you can try it in the playground (exact quote shown before you submit), then call it through one unified REST API — a typical run costs about $1.25.

TL;DR

· Task type: Omni Reference, provided by Alibaba

· Pricing: about $1.25 per typical run ($1 = 100 credits), pay-as-you-go

· Getting started: playground (exact quote before submit) → create an API key → one POST request

· Tunable parameters: 12 (see the table below)

· No platform watermark; commercial use allowed (per the Terms of Service)

What can you build with Wan 3.0 Prime?

Content automation: wire this model into your workflow and call it on demand through one unified API.

How do you call this API?

1. Sign up on NamiFusion and create a key on the API Keys page (signup bonus credits can only be used on selected models).

2. Submit a task with the cURL example below, passing your key as a Bearer token.

3. Poll the task status with the returned task_uuid and read the output URLs when completed.

Submit a task
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/alibaba/wan-3.0-prime/video" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "input": {
    "resolution": "1080P",
    "aspect_ratio": "adaptive",
    "duration": 5,
    "generate_audio": true,
    "enable_thinking": false
  }
}'

# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"

What parameters does it take?

prompt (Prompt): optional

images (Reference Images): optional

audios (Reference Audio): optional

first_frame (First Frame): optional

last_frame (Last Frame): optional

resolution (Resolution): optional, default "1080P"

aspect_ratio (Aspect Ratio): optional, default "adaptive"

duration (Duration (seconds)): optional, default 5

generate_audio (Generate Audio): optional, default true

enable_thinking (Thinking Mode): optional, default false

file (Reference Document): optional

link (Reference Web Page): optional

How do you write prompts for Wan 3.0 Prime?

1. Subject first: lead with the actor and action, then add environment, lighting and style ("an astronaut sprinting through a neon street in the rain, cinematic, shallow depth of field").

2. Concrete nouns beat stacked adjectives: "85mm portrait lens, golden-hour backlight" works far better than "very beautiful, ultra high quality".

3. Change one variable per iteration: pin the seed to reproduce a result, tweak a single phrase at a time so you know what actually moved the output.

4. State negatives positively: instead of "no blur", ask for "sharp focus, crisp details".

How much does it cost?

Pay-per-use in credits. A typical run costs about $1.25 ($1 = 100 credits); the exact price varies with parameters like resolution or duration, and the playground shows the exact quote before you submit.

Wan 3.0 Prime

One model for every input — text, images, video, audio, and documents — generating up to 30 seconds in a single shot.

Wan 3.0 Video is an all-in-one video generation model that upgrades duration, universal reference, and realism together. A single request produces up to 30 seconds of continuous video, giving narrative pacing room to breathe and letting complex camera language — continuous movement, one-take shots — play out fully. Beyond text, image, audio, and video, it is the first in the series to accept documents (doc, xls, ppt, pdf, md and more), turning office material directly into video.

🚀 Key Features

  • Native 30-second generation: Produce a complete story rather than a single shot. Smart duration lets the model recommend the right length from your prompt.
  • Omni reference across six modalities: Combine reference images, video, audio, documents, and web pages in one request, orchestrated by the prompt.
  • Documents as a source: doc, xls, ppt, pdf, txt, key, pages, numbers, and md become teaching decks, product demos, animated charts, and business reports.
  • Photorealistic world renderer: Faithful skin and facial detail with restrained, natural emotion — distinct faces rather than the same generic look, holding up even in crowd scenes.
  • Pixel-level consistency: Characters, props, sound, and spatial relationships stay aligned to your references — not "close enough", but faithfully reproduced.
  • Strict first/last frame mode: Pin the exact opening and closing frames when you need controlled transitions.
  • Native audio output: Generated video carries its own audio track; the switch is free either way.
  • Adaptive aspect ratio: Let the model pick the ratio from your intent and input media, or choose from five presets.

🛠️ Technical Specifications

ItemDetails
Model architectureAll-in-one omni-reference video generation model
Task typeVideo generation (text / image / video / audio / document to video)
PromptChinese and English, up to 5000 characters. Either prompt or reference media is required
Output FormatVideo with native audio
Resolution480P, 720P, 1080P (default 1080P)
Duration2–30 seconds (default 5); -1 enables smart duration. With video input, input plus output must not exceed 30 seconds
Aspect ratioadaptive (default), 16:9, 4:3, 1:1, 3:4, 9:16
Reference imagesUp to 10; JPEG, JPG, PNG (no alpha), BMP, WEBP; ≤20MB each; 240–8000px per side; ratio within 8:1
Reference videoUp to 5 clips, 1–15s each, 15s total; mp4, mov; <100MB each; 240–4096px per side
Reference audioUp to 5 clips, 1–15s each, 15s total; wav, mp3; <15MB each
DocumentsOne file, ≤100MB, ≤50 pages; docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md. Requires thinking mode
Web pagesOne public URL, no login required. Requires thinking mode; mutually exclusive with documents
First / last frameOne image each; mutually exclusive with omni-reference inputs
Thinking modeOff by default; required for documents, web pages, and complex image reasoning
SeedOptional; 0–2147483647 for reproducibility
Audio switchOn by default; pricing is unchanged either way
LatencyNo cold start

Sample Prompts

FAQ

Can I try it for free first?+

Signup bonus credits can only be used on selected models, and eligibility may change. The playground shows the exact quote before submission.

Can I use the output commercially? Is there a watermark?+

Yes, outputs are yours to use commercially, and NamiFusion adds no platform watermark — subject to the Terms of Service and applicable law.

What is the content policy?+

NamiFusion is an unfiltered engine: no arbitrary restrictions and NSFW-friendly, but illegal content and non-consensual use of real people are prohibited — see the Terms of Service.

How long does a Wan 3.0 Prime task take?+

It depends on parameters — images usually take seconds, videos tens of seconds to minutes. The API returns a task_uuid for polling.

What are the resolution and duration limits?+

See the parameter table above — each model lists its available resolution and duration options there and in the playground.

How to Use the Wan 3.0 Prime API: Omni Reference integration guide (2026) | NamiFusion Blog