seedance-2.0-mini/image-to-video-spicy

doubao/seedance-2.0-mini/image-to-video-spicy

Seedance 2.0 Mini Spicy Image to Video is ByteDance's faster, lower-cost image-to-video model for cinematic multi-shot videos. It turns reference images and optional text prompts into narrative sequences with AI camera control, consistent characters, 480P / 720P / 1080P / 4K output, 4-15s duration, and flexible aspect ratios. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Examples

seedance-2.0-mini/image-to-video-spicy Examples 1

Parameters

NameTypeDefaultConstraintsDescription
first_frame *Imageimage_upload—image/* · 0–1 itemsStart image URL to guide the video generation.
promptPrompttextarea—≤ 5000 charsDescribe the scene, action, camera movement, and mood for the video.
last_frameLast Frameimage_upload—image/* · 0–1 itemsLast frame image URL for video continuation.
aspect_ratioAspect Ratioselect—16:9 | 9:16 | 4:3 | 3:4 | 1:1 | 21:9The aspect ratio of the generated video. If not specified, adapts to the input image.
resolutionResolutionselect720p480p | 720p | 1080p | 4kThe output video resolution.
durationDurationslider54 ~ 15 · step 1The duration of the generated video in seconds (4-15s).
generate_audioGenerate Audiobooleantrue—Whether to generate native audio synchronized with the output video. Defaults to true.
seedSeednumber——The random seed to use for the generation. -1 means a random seed will be used.

Output fields

FieldTypeDescription
videosarray<string>Generated video URL(s).

API

Call this model through one unified REST API. Get a key on the API Keys page.

cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/doubao/seedance-2.0-mini/image-to-video-spicy" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "input": {
    "first_frame": [
      "https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/ae63692086b8.jpeg"
    ],
    "prompt": "The woman slowly turns toward the camera and begins walking forward through the rain as passing headlights streak across the frame. The camera starts with a medium shot, then gently dollies in and arcs slightly around her for a cinematic reveal. Moody, sensual atmosphere with neon reflections, drifting steam, subtle wind in her hair, and realistic motion. High-end film look, rich contrast, smooth camera movement, consistent character.",
    "last_frame": [
      "https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/50b4fab428d3.jpeg"
    ],
    "aspect_ratio": "16:9",
    "resolution": "720p",
    "duration": 6,
    "generate_audio": true,
    "seed": -1
  }
}'

# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"
Python
import time, requests

API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}

# 1) Submit
resp = requests.post(
    "https://www.namifusion.com/api/v1/marketplace/run/doubao/seedance-2.0-mini/image-to-video-spicy",
    headers=HEADERS,
    json={
        "input": {
            "first_frame": [
                "https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/ae63692086b8.jpeg"
            ],
            "prompt": "The woman slowly turns toward the camera and begins walking forward through the rain as passing headlights streak across the frame. The camera starts with a medium shot, then gently dollies in and arcs slightly around her for a cinematic reveal. Moody, sensual atmosphere with neon reflections, drifting steam, subtle wind in her hair, and realistic motion. High-end film look, rich contrast, smooth camera movement, consistent character.",
            "last_frame": [
                "https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/50b4fab428d3.jpeg"
            ],
            "aspect_ratio": "16:9",
            "resolution": "720p",
            "duration": 6,
            "generate_audio": True,
            "seed": -1
        }
    },
)
resp.raise_for_status()  # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()

# 2) Poll until a terminal state (completed / failed / cancelled).
#    This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
    if time.time() > deadline:
        raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
    time.sleep(3)
    poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
    poll.raise_for_status()
    task = poll.json()

print(task["status"], task.get("output"))
JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };

// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/doubao/seedance-2.0-mini/image-to-video-spicy", {
  method: "POST",
  headers: { ...HEADERS, "Content-Type": "application/json" },
  body: JSON.stringify({
    "input": {
      "first_frame": [
        "https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/ae63692086b8.jpeg"
      ],
      "prompt": "The woman slowly turns toward the camera and begins walking forward through the rain as passing headlights streak across the frame. The camera starts with a medium shot, then gently dollies in and arcs slightly around her for a cinematic reveal. Moody, sensual atmosphere with neon reflections, drifting steam, subtle wind in her hair, and realistic motion. High-end film look, rich contrast, smooth camera movement, consistent character.",
      "last_frame": [
        "https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/50b4fab428d3.jpeg"
      ],
      "aspect_ratio": "16:9",
      "resolution": "720p",
      "duration": 6,
      "generate_audio": true,
      "seed": -1
    }
  }),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();

// 2) Poll until a terminal state (completed / failed / cancelled).
//    This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
  if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
  await new Promise((r) => setTimeout(r, 3000));
  const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
  if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
  task = await poll.json();
}

console.log(task.status, task.output);

Documentation

doubao/seedance-2.0-mini/image-to-video-spicy

A faster, lower-cost image-to-video model for cinematic multi-shot storytelling with native audio.

doubao/seedance-2.0-mini/image-to-video-spicy is ByteDance’s lightweight image-to-video model designed for fast, affordable generation of narrative short-form videos. It transforms a reference image, optional Prompt, and optional ending-frame guidance into dynamic clips with camera-aware motion, character consistency, and flexible output settings from 480p to 4k.

This model is well suited for social content, creative prototyping, and cinematic short clips where speed and cost efficiency matter. Key advantages include no cold start, flexible Aspect Ratio support, native audio generation, and strong control over motion and camera language.

🚀 Key Features

  • Image-to-video generation: Turn a single reference Image into a motion-driven Video sequence.
  • Cinematic multi-shot output: Use Prompt guidance to describe action, mood, pacing, and camera behavior for more narrative results.
  • Last-frame guidance: Provide last_frame to influence the ending frame or continuation direction of the clip.
  • Native audio generation: Generate synchronized Audio together with the Video for more complete outputs.
  • Flexible output options: Supports 16:9, 9:16, 4:3, 3:4, 1:1, and 21:9 Aspect Ratio settings, plus 480p, 720p, 1080p, and 4k Resolution.
  • Fast and cost-efficient: Optimized for affordable production workloads with reliable performance and no cold starts.

🛠️ Technical Specifications

ItemDetails
Model architectureImage-to-Video model for cinematic short-form and multi-shot generation
InputStart Image URL, optional Prompt, optional last-frame Image URL
OutputVideo with optional native Audio
Output FormatVideo
Resolution480p, 720p(default), 1080p, 4k
Duration4–15 seconds(default: 5)
Aspect Ratio16:9, 9:16, 4:3, 3:4, 1:1, 21:9; adapts to the input Image if unspecified
AudioSupported; synchronized native Audio generation
SeedSupported; -1 uses a random Seed
Character consistencyMaintains subject continuity from the reference Image
Camera controlPrompt-based control for push-ins, pans, tracking, handheld motion, and more
LatencyOptimized for fast inference with no cold start

Sample Prompts

  1. A cinematic shot of the character slowly walking through a neon-lit street at night, soft rain falling, reflections on the wet pavement, gentle camera movement, atmospheric lighting, calm but dramatic mood.
  2. 0-2s: The product slowly rotates on a studio table under soft lighting. 2-5s: The camera slides sideways to reveal metallic edges and material detail. No subtitles.
  3. The parked car in the image pulls out and drives down the wet street. 0-3s: headlights flick on, the car eases forward. 3-6s: the camera pans to follow from a low angle, neon reflections on the asphalt. Ambient and engine sounds only, no music.

💰 Pricing

BillingResolutionPrice
Per second480p$0.06
Per second720p$0.12
Per second1080p$0.30
Per second4k$0.60
Per 5 seconds480p$0.30
Per 5 seconds720p$0.60
Per 5 seconds1080p$1.50
Per 5 seconds4k$3.00

Example Costs

Resolution4s5s10s15s
480p$0.24$0.30$0.60$0.90
720p$0.48$0.60$1.20$1.80
1080p$1.20$1.50$3.00$4.50
4k$2.40$3.00$6.00$9.00

💡 Best Use Cases

  • Social media video production: Create vertical, square, landscape, or ultrawide clips for platform-specific publishing.
  • E-commerce and product marketing: Animate product Images into polished showcase Videos with camera motion and sound.
  • Previsualization and concept testing: Rapidly explore scene direction, motion ideas, and shot design from a single reference Image.
  • Character and artwork animation: Turn illustrations, portraits, or concept art into short narrative Video sequences.

🔗 Related Models

  • ByteDance Seedance 2.0 Mini Image-to-Video: Standard image-to-video version for general-purpose generation.
  • ByteDance Seedance 2.0 Mini Text-to-Video: Generates Video directly from Prompt input without a reference Image.

Nami API usage

Send JSON { "input": { ... } } with an X-API-Key header. Each image parameter is an array containing one URL. Defaults below apply when a parameter is omitted; the playground and API examples use the example values shown below.

POST /api/v1/marketplace/run/doubao/seedance-2.0-mini/image-to-video-spicy

ParameterTypeRequiredDefaultConstraints
first_framestring[]yes—1 URL
promptstringno—max 5000 characters
last_framestring[]no—0–1 URL
aspect_ratiostringno—16:9, 9:16, 4:3, 3:4, 1:1, 21:9
resolutionstringno"720p"480p, 720p, 1080p, 4k
durationintegerno54–15; step 1
generate_audiobooleannotrue—
seedintegerno——
{
  "input": {
    "first_frame": [
      "https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/ae63692086b8.jpeg"
    ],
    "prompt": "The woman slowly turns toward the camera and begins walking forward through the rain as passing headlights streak across the frame. The camera starts with a medium shot, then gently dollies in and arcs slightly around her for a cinematic reveal. Moody, sensual atmosphere with neon reflections, drifting steam, subtle wind in her hair, and realistic motion. High-end film look, rich contrast, smooth camera movement, consistent character.",
    "last_frame": [
      "https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/50b4fab428d3.jpeg"
    ],
    "aspect_ratio": "16:9",
    "resolution": "720p",
    "duration": 6,
    "generate_audio": true,
    "seed": -1
  }
}

The create response contains task_uuid and task status. Poll using the same API key:

GET /api/v1/marketplace/run/tasks/{task_uuid}

On completion, read generated video URLs from output.videos.

{
  "status": "completed",
  "output": {
    "videos": [
      "https://cdn.example.com/output/video.mp4"
    ]
  }
}

This is a response excerpt; see the API tab for the full task structure.

List price is calculated from the selected resolution and duration; 1 USD = 100 credits.

ResolutionUSD / secondCredits / secondUSD / 5 seconds
480p$0.066$0.30
720p$0.1212$0.60
1080p$0.3030$1.50
4k$0.6060$3.00

Enabling or disabling audio does not change the list price.

Related models

xAI Grok Imagine Video v1.5 Image to Video
Image to VideoX Ai

xAI Grok Imagine Video v1.5 Image to Video

Animate one input image with a text prompt into a 1-15 second video at 480p or 720p.

from $0.850 / per run
vidu/q3/image-to-video-spicy
Image to VideoVidu

vidu/q3/image-to-video-spicy

Vidu Q3 Image-to-Video Spicy generates unlimited high-quality videos from images with smooth animations and diverse motion, optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.750 / per run
Vidu Q3 Image To Video
Image to VideoVidu

Vidu Q3 Image To Video

Vidu Q3 Image-to-Video turns text prompts into high-quality videos with exceptional visual fidelity and diverse motion. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.750 / per run
MiniMax H3 Spicy Image to Video
Image to VideoMinimax

MiniMax H3 Spicy Image to Video

Animate a first-frame image with optional prompt and last-frame guidance. Supports 480p, 540p, 768p and 1080p video, with durations from 3 to 15 seconds.

from $0.200 / per run
MiniMax H3 Image to Video
Image to VideoMinimax

MiniMax H3 Image to Video

MiniMax H3 Image to Video animates a first-frame image into a coherent 2K video, with natural-language motion instructions and optional last-frame control for consistent motion, scene continuity, and cinematic video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.650 / per run
ltx-2.5/image-to-video
Image to VideoLtx 2.5

ltx-2.5/image-to-video

LTX 2.5 Image-to-Video animates a first-frame image into high-fidelity synchronized audio-video content, with optional last-frame guidance and 720P / 1080P / 2K / 4K output for cinematic videos, social content, ads, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.500 / per run
ai/ltx-2.3-spicy/image-to-video-lora
Image to VideoLightricks

ai/ltx-2.3-spicy/image-to-video-lora

LTX 2.3 Spicy LoRA Image to Video API turns a reference image and prompt into expressive AI videos using selectable LoRA presets, optional LoRA strength overrides, duration, and resolution controls. Run fast REST inference on MaaS with no cold starts and affordable pricing.

from $0.150 / per run
ai/ltx-2.3-spicy/image-to-video
Image to VideoLightricks

ai/ltx-2.3-spicy/image-to-video

LTX 2.3 Spicy Image-to-Video API generates expressive AI videos from a reference image and prompt. Choose style presets, duration, and resolution with fast REST inference, no cold starts, and affordable pricing on MaaS.

from $0.100 / per run