vidu/q3/image-to-video-spicy
vidu/q3/image-to-video-spicy
Vidu Q3 Image-to-Video Spicy generates unlimited high-quality videos from images with smooth animations and diverse motion, optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Examples
Parameters
| Name | Type | Default | Constraints | Description |
|---|---|---|---|---|
| image *Image | image_upload | — | image/* · 1–1 items | The URL of the image to generate an image from. |
| promptPrompt | textarea | — | ≤ 5000 chars | The positive prompt for the generation. |
| resolutionResolution | select | 720p | 540p | 720p | 1080p | The resolution of the generated media. |
| durationDuration | slider | 5 | 1 ~ 16 · step 1 | The duration of the generated media in seconds. |
| movement_amplitudeMovement Amplitude | select | auto | auto | small | medium | large | The movement amplitude of objects in the frame. Defaults to auto, accepted value: auto small medium large. |
| generate_audioGenerate Audio | boolean | true | — | Whether to generate audio. |
| bgmBgm | boolean | true | — | The background music for generating the output. |
Output fields
| Field | Type | Description |
|---|---|---|
| videos | array<string> | Generated video URL(s). |
API
Call this model through one unified REST API. Get a key on the API Keys page.
cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/vidu/q3/image-to-video-spicy" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"prompt": "Animate the reference image with gentle natural movement and a slow cinematic camera push-in. Keep the subject, colors, and background consistent throughout the shot.",
"resolution": "720p",
"duration": 5,
"movement_amplitude": "auto",
"generate_audio": true,
"bgm": true
}
}'
# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"Python
import time, requests
API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
# 1) Submit
resp = requests.post(
"https://www.namifusion.com/api/v1/marketplace/run/vidu/q3/image-to-video-spicy",
headers=HEADERS,
json={
"input": {
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"prompt": "Animate the reference image with gentle natural movement and a slow cinematic camera push-in. Keep the subject, colors, and background consistent throughout the shot.",
"resolution": "720p",
"duration": 5,
"movement_amplitude": "auto",
"generate_audio": True,
"bgm": True
}
},
)
resp.raise_for_status() # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()
# 2) Poll until a terminal state (completed / failed / cancelled).
# This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
if time.time() > deadline:
raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
time.sleep(3)
poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
poll.raise_for_status()
task = poll.json()
print(task["status"], task.get("output"))JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };
// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/vidu/q3/image-to-video-spicy", {
method: "POST",
headers: { ...HEADERS, "Content-Type": "application/json" },
body: JSON.stringify({
"input": {
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"prompt": "Animate the reference image with gentle natural movement and a slow cinematic camera push-in. Keep the subject, colors, and background consistent throughout the shot.",
"resolution": "720p",
"duration": 5,
"movement_amplitude": "auto",
"generate_audio": true,
"bgm": true
}
}),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();
// 2) Poll until a terminal state (completed / failed / cancelled).
// This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
await new Promise((r) => setTimeout(r, 3000));
const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
task = await poll.json();
}
console.log(task.status, task.output);Documentation
vidu/q3/image-to-video-spicy
Turn a single image into vivid, high-quality video with smooth motion and production-ready scalability.
vidu/q3/image-to-video-spicy is an image-to-video model designed to transform still images into dynamic clips with expressive motion, natural transitions, and strong visual consistency. Optimized for scalable content generation, it combines flexible resolution options, up to 16 seconds of duration, optional audio generation, and no cold starts for reliable production use.
🚀 Key Features
- High-quality image-to-video generation: Converts a single image into polished video clips with stable aesthetics and smooth animation.
- Expressive motion control: Supports multiple motion intensity levels, from subtle movement to bold action, for a wide range of creative styles.
- Flexible resolution options: Generate at 540p, 720p, or 1080p depending on quality targets and budget.
- Duration control: Create clips from 1 to 16 seconds, making the model suitable for short-form content and promotional assets.
- Optional audio output: Generate synchronized Audio and include background music to reduce post-production effort.
- No cold start performance: Built for responsive, scalable workloads with consistent inference availability.
🛠️ Technical Specifications
| Item | Details |
|---|---|
| Model architecture | Image-to-video generation model |
| Input format | Image URL, Prompt |
| Output Format | Video, optional Audio |
| Resolution | 540p, 720p(default), 1080p |
| Duration | 1–16 seconds(default: 5) |
| Motion | auto, small, medium, large |
| Audio | Supported(enabled by default) |
| Background music | Supported(enabled by default) |
| Seed | Supported; -1 uses a random Seed |
| Latency | Not fixed in public specs; positioned for no cold start and stable inference performance |
| Typical applications | Marketing creatives, social content, product visuals, concept animation |
Sample Prompts
- Subtle head turn and gentle smile, soft wind moving the hair, slow camera push-in, cinematic lighting.
- Make the product slowly rotate with clean studio reflections and premium lighting, minimal background motion.
- Natural cloud drift and shifting sunlight, slight camera pan, dreamy atmosphere with smooth transitions.
💰 Pricing
| Resolution | Cost per second |
|---|---|
| 540p | $0.07 |
| 720p | $0.15 |
| 1080p | $0.16 |
Note: A base price of $0.35 is also provided in the source information. For practical cost estimation, the per-second pricing by resolution is the clearest reference.
💡 Best Use Cases
- E-commerce and product marketing: Animate product images into short promotional clips for storefronts, ads, and landing pages.
- Social media content production: Turn posters, illustrations, and hero images into engaging short-form video assets at scale.
- Creative storytelling and pitching: Add motion to concept art, character stills, or scene boards for more compelling visual presentations.
- Studio and media workflows: Produce large volumes of video variations efficiently for campaigns, publishing, or creative testing.
🔗 Related Models
- Vidu Q3 Image-to-Video: Standard tier with lower cost, suitable for budget-sensitive workloads.
- Vidu Q3 Text-to-Video: Generates video directly from text descriptions when no source image is available.
Related models
xAI Grok Imagine Video v1.5 Image to Video
Animate one input image with a text prompt into a 1-15 second video at 480p or 720p.
Vidu Q3 Image To Video
Vidu Q3 Image-to-Video turns text prompts into high-quality videos with exceptional visual fidelity and diverse motion. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

MiniMax H3 Spicy Image to Video
Animate a first-frame image with optional prompt and last-frame guidance. Supports 480p, 540p, 768p and 1080p video, with durations from 3 to 15 seconds.
MiniMax H3 Image to Video
MiniMax H3 Image to Video animates a first-frame image into a coherent 2K video, with natural-language motion instructions and optional last-frame control for consistent motion, scene continuity, and cinematic video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ltx-2.5/image-to-video
LTX 2.5 Image-to-Video animates a first-frame image into high-fidelity synchronized audio-video content, with optional last-frame guidance and 720P / 1080P / 2K / 4K output for cinematic videos, social content, ads, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ai/ltx-2.3-spicy/image-to-video-lora
LTX 2.3 Spicy LoRA Image to Video API turns a reference image and prompt into expressive AI videos using selectable LoRA presets, optional LoRA strength overrides, duration, and resolution controls. Run fast REST inference on MaaS with no cold starts and affordable pricing.
ai/ltx-2.3-spicy/image-to-video
LTX 2.3 Spicy Image-to-Video API generates expressive AI videos from a reference image and prompt. Choose style presets, duration, and resolution with fast REST inference, no cold starts, and affordable pricing on MaaS.
Kling Omni Video O3 Image-To-Video
Kling Omni Video O3 Image-to-Video transforms static images into dynamic cinematic videos using MVL (Multi-modal Visual Language) technology. Maintains subject consistency while adding natural motion, physics simulation, and seamless scene dynamics. Supports audio generation. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.