Omni ReferenceAlibaba

Wan 3.0 Omni Reference

alibaba/wan-3.0/video

alibaba/wan-3.0/video. Przeglądaj API modeli obrazu, wideo i zamiany twarzy. Wypróbuj model, a następnie zintegruj go przez wspólne REST API.

Powiązane modele

Wan 3.0 Omni Reference Powiązane modele 1

Pełne parametry, schemat wyjścia i aktualne ceny

Rynek API modeli AIJedna platforma do wszystkiegoNajpopularniejszyRegulaminPrzeglądaj API modeli obrazu, wideo i zamiany twarzy. Wypróbuj model, a następnie zintegruj go przez wspólne REST API.
promptPrompttextarea—≤ 5000 chars
imagesReference Imagesimage_upload—image/jpeg,image/png,image/bmp,image/webp · 0–10 items
audiosReference Audioaudio_upload_group—audio/wav,audio/mpeg · 0–5 items
first_frameFirst Frameimage_upload—image/jpeg,image/png,image/bmp,image/webp · 0–1 items
last_frameLast Frameimage_upload—image/jpeg,image/png,image/bmp,image/webp · 0–1 items
resolutionResolutionselect1080P480P | 720P | 1080P
aspect_ratioAspect Ratioselectadaptiveadaptive | 16:9 | 4:3 | 1:1 | 3:4 | 9:16
durationDuration (seconds)number5-1 ~ 30 · step 1
generate_audioGenerate Audiobooleantrue—
enable_thinkingThinking Modebooleanfalse—
fileReference Documentfile_upload—.docx,.doc,.xlsx,.xls,.pptx,.ppt,.pdf,.txt,.key,.pages,.numbers,.md
linkReference Web Pagetext——
seedSeednumber—0 ~ 2147483647 · step 1

Pełne parametry, schemat wyjścia i aktualne ceny

Dowiedz się więcejJedna platforma do wszystkiegoPrzeglądaj API modeli obrazu, wideo i zamiany twarzy. Wypróbuj model, a następnie zintegruj go przez wspólne REST API.
videosarray<string>

API

Przeglądaj API modeli obrazu, wideo i zamiany twarzy. Wypróbuj model, a następnie zintegruj go przez wspólne REST API.

cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/alibaba/wan-3.0/video" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "input": {
    "prompt": "A small sailboat glides across a calm lake at sunrise. Gentle ripples reflect the warm light as the camera slowly follows the boat in a smooth cinematic shot.",
    "resolution": "1080P",
    "aspect_ratio": "adaptive",
    "duration": 5,
    "generate_audio": true,
    "enable_thinking": false
  }
}'

# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"
Python
import time, requests

API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}

# 1) Submit
resp = requests.post(
    "https://www.namifusion.com/api/v1/marketplace/run/alibaba/wan-3.0/video",
    headers=HEADERS,
    json={
        "input": {
            "prompt": "A small sailboat glides across a calm lake at sunrise. Gentle ripples reflect the warm light as the camera slowly follows the boat in a smooth cinematic shot.",
            "resolution": "1080P",
            "aspect_ratio": "adaptive",
            "duration": 5,
            "generate_audio": True,
            "enable_thinking": False
        }
    },
)
resp.raise_for_status()  # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()

# 2) Poll until a terminal state (completed / failed / cancelled).
#    This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
    if time.time() > deadline:
        raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
    time.sleep(3)
    poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
    poll.raise_for_status()
    task = poll.json()

print(task["status"], task.get("output"))
JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };

// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/alibaba/wan-3.0/video", {
  method: "POST",
  headers: { ...HEADERS, "Content-Type": "application/json" },
  body: JSON.stringify({
    "input": {
      "prompt": "A small sailboat glides across a calm lake at sunrise. Gentle ripples reflect the warm light as the camera slowly follows the boat in a smooth cinematic shot.",
      "resolution": "1080P",
      "aspect_ratio": "adaptive",
      "duration": 5,
      "generate_audio": true,
      "enable_thinking": false
    }
  }),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();

// 2) Poll until a terminal state (completed / failed / cancelled).
//    This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
  if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
  await new Promise((r) => setTimeout(r, 3000));
  const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
  if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
  task = await poll.json();
}

console.log(task.status, task.output);

Dokumentacja API

Wan 3.0 Omni Reference

Turn images, video, audio, documents, and web pages into one video

Alibaba Wan 3.0 Omni Reference accepts five kinds of reference input at once and generates a single video of up to 30 seconds with a native audio track. It reproduces characters, props, and spatial relationships from your references at pixel-level fidelity, supports first/last-frame control for precise framing, and can decide the final length on its own. Ready-to-use REST inference API, no cold starts.

🚀 Key Features

  • Five reference modalities: Up to 10 reference images, 5 reference videos, and 5 audio clips, plus one document or one web link — combine them freely.
  • Pixel-level identity preservation: Characters, props, wardrobe, and spatial relationships stay consistent across new shots, with no extra training.
  • Native audio: Outputs video with sound by default — no separate dubbing pass.
  • First/last-frame control: Pin the opening frame, the closing frame, or both.
  • Model-decided duration: Hand the length to the model and let the content determine it (up to 30 seconds).
  • Thinking mode: Parses reference documents, web pages, and complex imagery before generating.

🛠️ Technical Specifications

ParameterDescription
Model ArchitectureAlibaba Wan 3.0
Reference ImagesUp to 10
Reference VideosUp to 5 (their duration is billed)
Reference AudioUp to 5
Reference Document / Web PageDocument ≤100MB and ≤50 pages, or one public URL; mutually exclusive, requires thinking mode
Output FormatVideo (MP4, audio included by default)
Resolution480P / 720P / 1080P (default 1080P)
Aspect Ratioadaptive / 16:9 / 4:3 / 1:1 / 3:4 / 9:16
Duration2–30 seconds (default 5), or -1 to let the model decide; with video input, input plus output must not exceed 30 seconds
Timeout300 seconds

💰 Pricing

ResolutionPricing formula
480P5 credits × total_duration
720P10 credits × total_duration
1080P20 credits × total_duration

total_duration = generated video duration + billed duration of reference video inputs (summed across all reference clips).

How model-decided duration is billed: setting duration to -1 lets the model choose the length, which cannot be known at request time, so it is billed at the 30-second maximum. For example, 1080P with model-decided duration and no reference video costs 600 credits per call. If you already know roughly how long the clip should be, passing an explicit number is cheaper.

💡 Best Use Cases

  • Character-consistent series: Reuse one reference set across many shots to keep a character and wardrobe identical.
  • Product photography to video: Feed product stills and footage as references to generate matching short-form clips.
  • Document or landing page to video: Turn on thinking mode and convert a spec sheet or web page into an explainer.
  • Social clips with sound: Native audio output removes the post-production dubbing step.
  • First/last-frame transitions: Pin start and end frames for precise cuts and hand-offs.

🔗 Related Models

  • alibaba/wan-2.6/reference-to-video: Replicates camera language and pacing from reference footage.
  • alibaba/wan-2.6/image-to-video: Single image to video at a lower cost.
  • alibaba/wan-2.6/text-to-video: Text-only video generation.