xAI Grok Imagine Video v1.5 Image to Video
x-ai/grok-imagine-video-v1.5/image-to-video
Animate one input image with a text prompt into a 1-15 second video at 480p or 720p.
Examples
Parameters
| Name | Type | Default | Constraints | Description |
|---|---|---|---|---|
| prompt *Prompt | textarea | — | ≤ 5000 chars | Text description of the desired motion, camera movement, and scene. |
| image *Image | image_upload | — | image/* · 0–1 items | Input image to animate. |
| durationDuration | slider | 6 | 1 ~ 15 · step 1 | Output video duration in seconds. |
| resolutionResolution | select | 720p | 720p | 480p | Output video resolution. |
Output fields
| Field | Type | Description |
|---|---|---|
| videos | array<string> | Generated video URL(s). |
API
Call this model through one unified REST API. Get a key on the API Keys page.
cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/x-ai/grok-imagine-video-v1.5/image-to-video" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"prompt": "Animate the scene with realistic subject motion, a slow cinematic camera push-in, subtle environmental movement, and stable details.",
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"duration": 6,
"resolution": "720p"
}
}'
# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"Python
import time, requests
API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
# 1) Submit
resp = requests.post(
"https://www.namifusion.com/api/v1/marketplace/run/x-ai/grok-imagine-video-v1.5/image-to-video",
headers=HEADERS,
json={
"input": {
"prompt": "Animate the scene with realistic subject motion, a slow cinematic camera push-in, subtle environmental movement, and stable details.",
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"duration": 6,
"resolution": "720p"
}
},
)
resp.raise_for_status() # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()
# 2) Poll until a terminal state (completed / failed / cancelled).
# This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
if time.time() > deadline:
raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
time.sleep(3)
poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
poll.raise_for_status()
task = poll.json()
print(task["status"], task.get("output"))JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };
// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/x-ai/grok-imagine-video-v1.5/image-to-video", {
method: "POST",
headers: { ...HEADERS, "Content-Type": "application/json" },
body: JSON.stringify({
"input": {
"prompt": "Animate the scene with realistic subject motion, a slow cinematic camera push-in, subtle environmental movement, and stable details.",
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"duration": 6,
"resolution": "720p"
}
}),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();
// 2) Poll until a terminal state (completed / failed / cancelled).
// This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
await new Promise((r) => setTimeout(r, 3000));
const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
task = await poll.json();
}
console.log(task.status, task.output);Documentation
xAI Grok Imagine Video v1.5 Image to Video
Turn a single image into short, stylized videos with fast iteration and predictable cost.
🚀 Key Features
- Image-guided video generation: Start from a single reference image and animate it into a short video clip.
- Prompt-driven motion control: Use Prompt instructions to define subject movement, camera behavior, atmosphere, and scene progression.
- Optimized for short-form content: Well suited for 1–15 second clips where fast iteration and clean control matter.
- Flexible quality settings: Choose between 480p for lower-cost testing and 720p for higher-quality output.
- Built for creative and marketing workflows: Effective for character animation, product videos, social content, concept visualization, and lightweight storytelling.
🛠️ Technical Specifications
| Item | Specification |
|---|---|
| Model Name | xAI Grok Imagine Video v1.5 Image to Video |
| Task Type | Image-to-Video |
| Model Architecture | Reference-image and Prompt-based short video generation model |
| Input | Image, Prompt |
| Output Format | Video |
| Resolution | 480p, 720p |
| Default Resolution | 720p |
| Duration | 1–15 seconds |
| Default Duration | 6 seconds |
| Motion Control | Prompt-based control over subject motion, camera movement, and scene evolution |
| Output Style | Stylized short videos, cinematic concept clips, lightweight motion storytelling |
Sample Prompts
A cinematic push-in shot as the subject slowly turns toward the camera, soft natural motion, subtle background movement, realistic lighting, polished commercial styleA premium product on a clean tabletop, slow orbit camera movement, soft reflections, minimal background, elegant advertising lookA character standing in a neon-lit street, slight hair and clothing motion, gentle forward camera move, atmospheric futuristic mood
💰 Pricing
| Resolution | Price per Second | Input Image Fee | 5s Example |
|---|---|---|---|
| 480p | $0.08 | $0.01 / image | $0.41 |
| 720p | $0.14 | $0.01 / image | $0.71 |
Billing Rules
| Rule | Details |
|---|---|
| Pricing model | Linear pricing based on output duration |
| Duration billing | Rounded up to the next whole second |
| Minimum billed Duration | 1 second |
| Maximum billed Duration | 15 seconds |
| Additional charge | Each request includes a fixed $0.01 input image fee |
💡 Best Use Cases
- Social media content: Animate still visuals into short promotional clips for posts, reels, and lightweight campaigns.
- E-commerce and product marketing: Turn product images into motion-driven demos for ads, landing pages, and creative testing.
- Concept visualization: Explore camera direction, motion ideas, and scene atmosphere from a single static frame.
- Creative prototyping: Quickly test Prompt-based motion concepts for characters, environments, and branded storytelling.
🔗 Related Models
- xAI text-to-image workflows: Useful when you want to generate the source image before animating it.
- Other image-to-video workflows: Worth comparing when you need different trade-offs in quality, speed, or motion control.
Related models
vidu/q3/image-to-video-spicy
Vidu Q3 Image-to-Video Spicy generates unlimited high-quality videos from images with smooth animations and diverse motion, optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Vidu Q3 Image To Video
Vidu Q3 Image-to-Video turns text prompts into high-quality videos with exceptional visual fidelity and diverse motion. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

MiniMax H3 Spicy Image to Video
Animate a first-frame image with optional prompt and last-frame guidance. Supports 480p, 540p, 768p and 1080p video, with durations from 3 to 15 seconds.
MiniMax H3 Image to Video
MiniMax H3 Image to Video animates a first-frame image into a coherent 2K video, with natural-language motion instructions and optional last-frame control for consistent motion, scene continuity, and cinematic video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ltx-2.5/image-to-video
LTX 2.5 Image-to-Video animates a first-frame image into high-fidelity synchronized audio-video content, with optional last-frame guidance and 720P / 1080P / 2K / 4K output for cinematic videos, social content, ads, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ai/ltx-2.3-spicy/image-to-video-lora
LTX 2.3 Spicy LoRA Image to Video API turns a reference image and prompt into expressive AI videos using selectable LoRA presets, optional LoRA strength overrides, duration, and resolution controls. Run fast REST inference on MaaS with no cold starts and affordable pricing.
ai/ltx-2.3-spicy/image-to-video
LTX 2.3 Spicy Image-to-Video API generates expressive AI videos from a reference image and prompt. Choose style presets, duration, and resolution with fast REST inference, no cold starts, and affordable pricing on MaaS.
Kling Omni Video O3 Image-To-Video
Kling Omni Video O3 Image-to-Video transforms static images into dynamic cinematic videos using MVL (Multi-modal Visual Language) technology. Maintains subject consistency while adding natural motion, physics simulation, and seamless scene dynamics. Supports audio generation. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.