alibaba/wan-2.6/image-to-video-spicy

alibaba/wan-2.6/image-to-video-spicy

WAN 2.6 Spicy converts images into unlimited high-quality videos with smooth animations optimized for scalable content generation. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Examples

alibaba/wan-2.6/image-to-video-spicy Examples 1

Parameters

NameTypeDefaultConstraintsDescription
image_url *Imageimage_upload—1–1 itemsThe image for generating the output.
audio_urlAudioc URLaudio_upload——Audio URL to guide generation (optional).
promptPrompttextarea—≤ 1500 charsThe positive prompt for the generation.
negative_promptNegative Prompttextarea—≤ 1500 charsThe negative prompt for the generation.
resolutionResolutionselect720p720p | 1080pThe resolution of the generated media.
durationDurationselect55 | 10 | 15The duration of the generated media in seconds.
shot_typeShot Typeselectsinglesingle | multiThe type of shots to generate.
prompt_extendEnable Prompt Expansionbooleantrue—If set to true, the prompt optimizer will be enabled.
seedSeednumber—≥ 0The random seed to use for the generation.

Output fields

FieldTypeDescription
videosarray<string>Generated video URL(s).

API

Call this model through one unified REST API. Get a key on the API Keys page.

cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/alibaba/wan-2.6/image-to-video-spicy" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "input": {
    "image_url": [
      "https://assets-public.namifusion.com/uploads/images/2026-03-05/559f2f35a8de.png"
    ],
    "prompt": "Create a dynamic video showcasing the transition from day to night in a busy city.",
    "negative_prompt": "Avoid using any blurred or low-resolution scenes.",
    "resolution": "1080P",
    "duration": 10,
    "shot_type": "multi",
    "prompt_extend": true,
    "seed": -1
  }
}'

# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"
Python
import time, requests

API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}

# 1) Submit
resp = requests.post(
    "https://www.namifusion.com/api/v1/marketplace/run/alibaba/wan-2.6/image-to-video-spicy",
    headers=HEADERS,
    json={
        "input": {
            "image_url": [
                "https://assets-public.namifusion.com/uploads/images/2026-03-05/559f2f35a8de.png"
            ],
            "prompt": "Create a dynamic video showcasing the transition from day to night in a busy city.",
            "negative_prompt": "Avoid using any blurred or low-resolution scenes.",
            "resolution": "1080P",
            "duration": 10,
            "shot_type": "multi",
            "prompt_extend": True,
            "seed": -1
        }
    },
)
resp.raise_for_status()  # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()

# 2) Poll until a terminal state (completed / failed / cancelled).
#    This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
    if time.time() > deadline:
        raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
    time.sleep(3)
    poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
    poll.raise_for_status()
    task = poll.json()

print(task["status"], task.get("output"))
JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };

// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/alibaba/wan-2.6/image-to-video-spicy", {
  method: "POST",
  headers: { ...HEADERS, "Content-Type": "application/json" },
  body: JSON.stringify({
    "input": {
      "image_url": [
        "https://assets-public.namifusion.com/uploads/images/2026-03-05/559f2f35a8de.png"
      ],
      "prompt": "Create a dynamic video showcasing the transition from day to night in a busy city.",
      "negative_prompt": "Avoid using any blurred or low-resolution scenes.",
      "resolution": "1080P",
      "duration": 10,
      "shot_type": "multi",
      "prompt_extend": true,
      "seed": -1
    }
  }),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();

// 2) Poll until a terminal state (completed / failed / cancelled).
//    This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
  if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
  await new Promise((r) => setTimeout(r, 3000));
  const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
  if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
  task = await poll.json();
}

console.log(task.status, task.output);

Documentation

alibaba/wan-2.6/image-to-video-spicy

Turn still images into expressive, high-quality videos with smooth cinematic motion.

alibaba/wan-2.6/image-to-video-spicy is an image-to-video model designed for scalable content production. It transforms a single image into dynamic video clips with natural transitions, vivid motion, and strong visual impact, while offering no cold starts and support for both 720p and native 1080p output.

🚀 Key Features

  • High-quality image-to-video generation: Animate still images into polished video clips with smooth, temporally coherent motion.
  • Native 720p and 1080p output: Choose the resolution that best fits your quality and budget requirements.
  • Flexible clip durations: Generate 5s, 10s, or 15s videos for short-form content, ads, and storytelling.
  • Single-shot or multi-shot modes: Use shot_type to create either stable single-shot animations or more dynamic multi-shot sequences.
  • Prompt expansion support: Enable automatic prompt optimization to enrich descriptions and improve generation quality.
  • No cold starts for production workflows: Built for responsive, scalable generation with consistent performance.

🛠️ Technical Specifications

ItemDetails
Model architectureImage-to-Video
InputImage URL, Prompt, Negative Prompt, optional Audio URL
Output FormatVideo
Resolution720p(default), 1080p
Duration5s(default), 10s, 15s
Shot typesingle, multi
Prompt expansionSupported(disabled by default)
SeedSupported; -1 uses a random Seed
Audio guidanceOptional Audio URL can be used to guide generation
Visual characteristicsSmooth animation, rich color, natural transitions, expressive motion
LatencyNo fixed latency published; positioned for no cold starts and stable inference

Sample Prompts

  • A cinematic slow push-in on a portrait, soft wind moving the hair, warm golden-hour lighting, natural motion, rich color contrast
  • A product shot of a perfume bottle on reflective glass, subtle camera orbit, sparkling highlights, premium advertising style
  • An illustrated fantasy city at night, drifting fog, glowing windows, gentle parallax, atmospheric storytelling

💰 Pricing

Resolution5s10s15s
720p$0.50$1.00$1.50
1080p$0.75$1.50$2.25

Billing note: 1080p is priced at a 1.5× multiplier over the 720p rate.

💡 Best Use Cases

  • Social media content: Turn cover art, illustrations, or product images into short-form videos for Reels, Shorts, and Stories.
  • Marketing and advertising: Produce scalable video creatives from product visuals for campaigns and performance ads.
  • Creative storytelling: Extend concept art, portraits, or scene images into cinematic clips with mood and motion.
  • Music and art projects: Pair visuals with optional audio guidance to create expressive, rhythm-aware video content.

🔗 Related Models

  • Wan 2.6 Image-to-Video Pro: Higher-tier version with 4K support for premium output needs.
  • Wan 2.6 Image-to-Video: Standard image-to-video variant.
  • Wan 2.6 Text-to-Video: Designed for generating videos directly from text prompts.

Best Practices

  • Use clear, well-framed images with a strong subject for more stable and visually convincing motion.
  • Write detailed Prompt descriptions covering motion, mood, camera movement, and style.
  • Use Negative Prompt to exclude unwanted elements, artifacts, or stylistic outcomes.
  • Enable prompt expansion when you want the model to automatically enrich sparse descriptions.

Related models

xAI Grok Imagine Video v1.5 Image to Video
Image to VideoX Ai

xAI Grok Imagine Video v1.5 Image to Video

Animate one input image with a text prompt into a 1-15 second video at 480p or 720p.

from $0.850 / per run
vidu/q3/image-to-video-spicy
Image to VideoVidu

vidu/q3/image-to-video-spicy

Vidu Q3 Image-to-Video Spicy generates unlimited high-quality videos from images with smooth animations and diverse motion, optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.750 / per run
Vidu Q3 Image To Video
Image to VideoVidu

Vidu Q3 Image To Video

Vidu Q3 Image-to-Video turns text prompts into high-quality videos with exceptional visual fidelity and diverse motion. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.750 / per run
MiniMax H3 Spicy Image to Video
Image to VideoMinimax

MiniMax H3 Spicy Image to Video

Animate a first-frame image with optional prompt and last-frame guidance. Supports 480p, 540p, 768p and 1080p video, with durations from 3 to 15 seconds.

from $0.200 / per run
MiniMax H3 Image to Video
Image to VideoMinimax

MiniMax H3 Image to Video

MiniMax H3 Image to Video animates a first-frame image into a coherent 2K video, with natural-language motion instructions and optional last-frame control for consistent motion, scene continuity, and cinematic video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.650 / per run
ltx-2.5/image-to-video
Image to VideoLtx 2.5

ltx-2.5/image-to-video

LTX 2.5 Image-to-Video animates a first-frame image into high-fidelity synchronized audio-video content, with optional last-frame guidance and 720P / 1080P / 2K / 4K output for cinematic videos, social content, ads, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.500 / per run
ai/ltx-2.3-spicy/image-to-video-lora
Image to VideoLightricks

ai/ltx-2.3-spicy/image-to-video-lora

LTX 2.3 Spicy LoRA Image to Video API turns a reference image and prompt into expressive AI videos using selectable LoRA presets, optional LoRA strength overrides, duration, and resolution controls. Run fast REST inference on MaaS with no cold starts and affordable pricing.

from $0.150 / per run
ai/ltx-2.3-spicy/image-to-video
Image to VideoLightricks

ai/ltx-2.3-spicy/image-to-video

LTX 2.3 Spicy Image-to-Video API generates expressive AI videos from a reference image and prompt. Choose style presets, duration, and resolution with fast REST inference, no cold starts, and affordable pricing on MaaS.

from $0.100 / per run