seedance-2.0-mini/image-to-video-spicy
doubao/seedance-2.0-mini/image-to-video-spicy
Seedance 2.0 Mini Spicy Image to Video is ByteDance's faster, lower-cost image-to-video model for cinematic multi-shot videos. It turns reference images and optional text prompts into narrative sequences with AI camera control, consistent characters, 480P / 720P / 1080P / 4K output, 4-15s duration, and flexible aspect ratios. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Examples
Parameters
| Name | Type | Default | Constraints | Description |
|---|---|---|---|---|
| first_frame *Image | image_upload | — | image/* · 0–1 items | Start image URL to guide the video generation. |
| promptPrompt | textarea | — | ≤ 5000 chars | Describe the scene, action, camera movement, and mood for the video. |
| last_frameLast Frame | image_upload | — | image/* · 0–1 items | Last frame image URL for video continuation. |
| aspect_ratioAspect Ratio | select | — | 16:9 | 9:16 | 4:3 | 3:4 | 1:1 | 21:9 | The aspect ratio of the generated video. If not specified, adapts to the input image. |
| resolutionResolution | select | 720p | 480p | 720p | 1080p | 4k | The output video resolution. |
| durationDuration | slider | 5 | 4 ~ 15 · step 1 | The duration of the generated video in seconds (4-15s). |
| generate_audioGenerate Audio | boolean | true | — | Whether to generate native audio synchronized with the output video. Defaults to true. |
| seedSeed | number | — | — | The random seed to use for the generation. -1 means a random seed will be used. |
Output fields
| Field | Type | Description |
|---|---|---|
| videos | array<string> | Generated video URL(s). |
API
Call this model through one unified REST API. Get a key on the API Keys page.
cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/doubao/seedance-2.0-mini/image-to-video-spicy" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"first_frame": [
"https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/ae63692086b8.jpeg"
],
"prompt": "The woman slowly turns toward the camera and begins walking forward through the rain as passing headlights streak across the frame. The camera starts with a medium shot, then gently dollies in and arcs slightly around her for a cinematic reveal. Moody, sensual atmosphere with neon reflections, drifting steam, subtle wind in her hair, and realistic motion. High-end film look, rich contrast, smooth camera movement, consistent character.",
"last_frame": [
"https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/50b4fab428d3.jpeg"
],
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 6,
"generate_audio": true,
"seed": -1
}
}'
# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"Python
import time, requests
API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
# 1) Submit
resp = requests.post(
"https://www.namifusion.com/api/v1/marketplace/run/doubao/seedance-2.0-mini/image-to-video-spicy",
headers=HEADERS,
json={
"input": {
"first_frame": [
"https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/ae63692086b8.jpeg"
],
"prompt": "The woman slowly turns toward the camera and begins walking forward through the rain as passing headlights streak across the frame. The camera starts with a medium shot, then gently dollies in and arcs slightly around her for a cinematic reveal. Moody, sensual atmosphere with neon reflections, drifting steam, subtle wind in her hair, and realistic motion. High-end film look, rich contrast, smooth camera movement, consistent character.",
"last_frame": [
"https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/50b4fab428d3.jpeg"
],
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 6,
"generate_audio": True,
"seed": -1
}
},
)
resp.raise_for_status() # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()
# 2) Poll until a terminal state (completed / failed / cancelled).
# This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
if time.time() > deadline:
raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
time.sleep(3)
poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
poll.raise_for_status()
task = poll.json()
print(task["status"], task.get("output"))JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };
// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/doubao/seedance-2.0-mini/image-to-video-spicy", {
method: "POST",
headers: { ...HEADERS, "Content-Type": "application/json" },
body: JSON.stringify({
"input": {
"first_frame": [
"https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/ae63692086b8.jpeg"
],
"prompt": "The woman slowly turns toward the camera and begins walking forward through the rain as passing headlights streak across the frame. The camera starts with a medium shot, then gently dollies in and arcs slightly around her for a cinematic reveal. Moody, sensual atmosphere with neon reflections, drifting steam, subtle wind in her hair, and realistic motion. High-end film look, rich contrast, smooth camera movement, consistent character.",
"last_frame": [
"https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/50b4fab428d3.jpeg"
],
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 6,
"generate_audio": true,
"seed": -1
}
}),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();
// 2) Poll until a terminal state (completed / failed / cancelled).
// This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
await new Promise((r) => setTimeout(r, 3000));
const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
task = await poll.json();
}
console.log(task.status, task.output);Documentation
doubao/seedance-2.0-mini/image-to-video-spicy
A faster, lower-cost image-to-video model for cinematic multi-shot storytelling with native audio.
doubao/seedance-2.0-mini/image-to-video-spicy is ByteDance’s lightweight image-to-video model designed for fast, affordable generation of narrative short-form videos. It transforms a reference image, optional Prompt, and optional ending-frame guidance into dynamic clips with camera-aware motion, character consistency, and flexible output settings from 480p to 4k.
This model is well suited for social content, creative prototyping, and cinematic short clips where speed and cost efficiency matter. Key advantages include no cold start, flexible Aspect Ratio support, native audio generation, and strong control over motion and camera language.
🚀 Key Features
- Image-to-video generation: Turn a single reference Image into a motion-driven Video sequence.
- Cinematic multi-shot output: Use Prompt guidance to describe action, mood, pacing, and camera behavior for more narrative results.
- Last-frame guidance: Provide
last_frameto influence the ending frame or continuation direction of the clip. - Native audio generation: Generate synchronized Audio together with the Video for more complete outputs.
- Flexible output options: Supports 16:9, 9:16, 4:3, 3:4, 1:1, and 21:9 Aspect Ratio settings, plus 480p, 720p, 1080p, and 4k Resolution.
- Fast and cost-efficient: Optimized for affordable production workloads with reliable performance and no cold starts.
🛠️ Technical Specifications
| Item | Details |
|---|---|
| Model architecture | Image-to-Video model for cinematic short-form and multi-shot generation |
| Input | Start Image URL, optional Prompt, optional last-frame Image URL |
| Output | Video with optional native Audio |
| Output Format | Video |
| Resolution | 480p, 720p(default), 1080p, 4k |
| Duration | 4–15 seconds(default: 5) |
| Aspect Ratio | 16:9, 9:16, 4:3, 3:4, 1:1, 21:9; adapts to the input Image if unspecified |
| Audio | Supported; synchronized native Audio generation |
| Seed | Supported; -1 uses a random Seed |
| Character consistency | Maintains subject continuity from the reference Image |
| Camera control | Prompt-based control for push-ins, pans, tracking, handheld motion, and more |
| Latency | Optimized for fast inference with no cold start |
Sample Prompts
A cinematic shot of the character slowly walking through a neon-lit street at night, soft rain falling, reflections on the wet pavement, gentle camera movement, atmospheric lighting, calm but dramatic mood.0-2s: The product slowly rotates on a studio table under soft lighting. 2-5s: The camera slides sideways to reveal metallic edges and material detail. No subtitles.The parked car in the image pulls out and drives down the wet street. 0-3s: headlights flick on, the car eases forward. 3-6s: the camera pans to follow from a low angle, neon reflections on the asphalt. Ambient and engine sounds only, no music.
💰 Pricing
| Billing | Resolution | Price |
|---|---|---|
| Per second | 480p | $0.06 |
| Per second | 720p | $0.12 |
| Per second | 1080p | $0.30 |
| Per second | 4k | $0.60 |
| Per 5 seconds | 480p | $0.30 |
| Per 5 seconds | 720p | $0.60 |
| Per 5 seconds | 1080p | $1.50 |
| Per 5 seconds | 4k | $3.00 |
Example Costs
| Resolution | 4s | 5s | 10s | 15s |
|---|---|---|---|---|
| 480p | $0.24 | $0.30 | $0.60 | $0.90 |
| 720p | $0.48 | $0.60 | $1.20 | $1.80 |
| 1080p | $1.20 | $1.50 | $3.00 | $4.50 |
| 4k | $2.40 | $3.00 | $6.00 | $9.00 |
💡 Best Use Cases
- Social media video production: Create vertical, square, landscape, or ultrawide clips for platform-specific publishing.
- E-commerce and product marketing: Animate product Images into polished showcase Videos with camera motion and sound.
- Previsualization and concept testing: Rapidly explore scene direction, motion ideas, and shot design from a single reference Image.
- Character and artwork animation: Turn illustrations, portraits, or concept art into short narrative Video sequences.
🔗 Related Models
- ByteDance Seedance 2.0 Mini Image-to-Video: Standard image-to-video version for general-purpose generation.
- ByteDance Seedance 2.0 Mini Text-to-Video: Generates Video directly from Prompt input without a reference Image.
Nami API usage
Send JSON { "input": { ... } } with an X-API-Key header. Each image parameter is an array containing one URL. Defaults below apply when a parameter is omitted; the playground and API examples use the example values shown below.
POST /api/v1/marketplace/run/doubao/seedance-2.0-mini/image-to-video-spicy
| Parameter | Type | Required | Default | Constraints |
|---|---|---|---|---|
first_frame | string[] | yes | — | 1 URL |
prompt | string | no | — | max 5000 characters |
last_frame | string[] | no | — | 0–1 URL |
aspect_ratio | string | no | — | 16:9, 9:16, 4:3, 3:4, 1:1, 21:9 |
resolution | string | no | "720p" | 480p, 720p, 1080p, 4k |
duration | integer | no | 5 | 4–15; step 1 |
generate_audio | boolean | no | true | — |
seed | integer | no | — | — |
{
"input": {
"first_frame": [
"https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/ae63692086b8.jpeg"
],
"prompt": "The woman slowly turns toward the camera and begins walking forward through the rain as passing headlights streak across the frame. The camera starts with a medium shot, then gently dollies in and arcs slightly around her for a cinematic reveal. Moody, sensual atmosphere with neon reflections, drifting steam, subtle wind in her hair, and realistic motion. High-end film look, rich contrast, smooth camera movement, consistent character.",
"last_frame": [
"https://d1s3annbwribom.cloudfront.net/marketplace/examples/2026/09/12/50b4fab428d3.jpeg"
],
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 6,
"generate_audio": true,
"seed": -1
}
}
The create response contains task_uuid and task status. Poll using the same API key:
GET /api/v1/marketplace/run/tasks/{task_uuid}
On completion, read generated video URLs from output.videos.
{
"status": "completed",
"output": {
"videos": [
"https://cdn.example.com/output/video.mp4"
]
}
}
This is a response excerpt; see the API tab for the full task structure.
List price is calculated from the selected resolution and duration; 1 USD = 100 credits.
| Resolution | USD / second | Credits / second | USD / 5 seconds |
|---|---|---|---|
| 480p | $0.06 | 6 | $0.30 |
| 720p | $0.12 | 12 | $0.60 |
| 1080p | $0.30 | 30 | $1.50 |
| 4k | $0.60 | 60 | $3.00 |
Enabling or disabling audio does not change the list price.
Related models
xAI Grok Imagine Video v1.5 Image to Video
Animate one input image with a text prompt into a 1-15 second video at 480p or 720p.
vidu/q3/image-to-video-spicy
Vidu Q3 Image-to-Video Spicy generates unlimited high-quality videos from images with smooth animations and diverse motion, optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Vidu Q3 Image To Video
Vidu Q3 Image-to-Video turns text prompts into high-quality videos with exceptional visual fidelity and diverse motion. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

MiniMax H3 Spicy Image to Video
Animate a first-frame image with optional prompt and last-frame guidance. Supports 480p, 540p, 768p and 1080p video, with durations from 3 to 15 seconds.
MiniMax H3 Image to Video
MiniMax H3 Image to Video animates a first-frame image into a coherent 2K video, with natural-language motion instructions and optional last-frame control for consistent motion, scene continuity, and cinematic video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ltx-2.5/image-to-video
LTX 2.5 Image-to-Video animates a first-frame image into high-fidelity synchronized audio-video content, with optional last-frame guidance and 720P / 1080P / 2K / 4K output for cinematic videos, social content, ads, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ai/ltx-2.3-spicy/image-to-video-lora
LTX 2.3 Spicy LoRA Image to Video API turns a reference image and prompt into expressive AI videos using selectable LoRA presets, optional LoRA strength overrides, duration, and resolution controls. Run fast REST inference on MaaS with no cold starts and affordable pricing.
ai/ltx-2.3-spicy/image-to-video
LTX 2.3 Spicy Image-to-Video API generates expressive AI videos from a reference image and prompt. Choose style presets, duration, and resolution with fast REST inference, no cold starts, and affordable pricing on MaaS.