xAI Grok Imagine Video v1.5 Image to Video

x-ai/grok-imagine-video-v1.5/image-to-video

Animate one input image with a text prompt into a 1-15 second video at 480p or 720p.

Examples

xAI Grok Imagine Video v1.5 Image to Video Examples 1

Parameters

NameTypeDefaultConstraintsDescription
prompt *Prompttextarea—≤ 5000 charsText description of the desired motion, camera movement, and scene.
image *Imageimage_upload—image/* · 0–1 itemsInput image to animate.
durationDurationslider61 ~ 15 · step 1Output video duration in seconds.
resolutionResolutionselect720p720p | 480pOutput video resolution.

Output fields

FieldTypeDescription
videosarray<string>Generated video URL(s).

API

Call this model through one unified REST API. Get a key on the API Keys page.

cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/x-ai/grok-imagine-video-v1.5/image-to-video" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "input": {
    "prompt": "Animate the scene with realistic subject motion, a slow cinematic camera push-in, subtle environmental movement, and stable details.",
    "image": [
      "https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
    ],
    "duration": 6,
    "resolution": "720p"
  }
}'

# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"
Python
import time, requests

API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}

# 1) Submit
resp = requests.post(
    "https://www.namifusion.com/api/v1/marketplace/run/x-ai/grok-imagine-video-v1.5/image-to-video",
    headers=HEADERS,
    json={
        "input": {
            "prompt": "Animate the scene with realistic subject motion, a slow cinematic camera push-in, subtle environmental movement, and stable details.",
            "image": [
                "https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
            ],
            "duration": 6,
            "resolution": "720p"
        }
    },
)
resp.raise_for_status()  # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()

# 2) Poll until a terminal state (completed / failed / cancelled).
#    This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
    if time.time() > deadline:
        raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
    time.sleep(3)
    poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
    poll.raise_for_status()
    task = poll.json()

print(task["status"], task.get("output"))
JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };

// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/x-ai/grok-imagine-video-v1.5/image-to-video", {
  method: "POST",
  headers: { ...HEADERS, "Content-Type": "application/json" },
  body: JSON.stringify({
    "input": {
      "prompt": "Animate the scene with realistic subject motion, a slow cinematic camera push-in, subtle environmental movement, and stable details.",
      "image": [
        "https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
      ],
      "duration": 6,
      "resolution": "720p"
    }
  }),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();

// 2) Poll until a terminal state (completed / failed / cancelled).
//    This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
  if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
  await new Promise((r) => setTimeout(r, 3000));
  const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
  if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
  task = await poll.json();
}

console.log(task.status, task.output);

Documentation

xAI Grok Imagine Video v1.5 Image to Video

Turn a single image into short, stylized videos with fast iteration and predictable cost.

🚀 Key Features

  • Image-guided video generation: Start from a single reference image and animate it into a short video clip.
  • Prompt-driven motion control: Use Prompt instructions to define subject movement, camera behavior, atmosphere, and scene progression.
  • Optimized for short-form content: Well suited for 1–15 second clips where fast iteration and clean control matter.
  • Flexible quality settings: Choose between 480p for lower-cost testing and 720p for higher-quality output.
  • Built for creative and marketing workflows: Effective for character animation, product videos, social content, concept visualization, and lightweight storytelling.

🛠️ Technical Specifications

ItemSpecification
Model NamexAI Grok Imagine Video v1.5 Image to Video
Task TypeImage-to-Video
Model ArchitectureReference-image and Prompt-based short video generation model
InputImage, Prompt
Output FormatVideo
Resolution480p, 720p
Default Resolution720p
Duration1–15 seconds
Default Duration6 seconds
Motion ControlPrompt-based control over subject motion, camera movement, and scene evolution
Output StyleStylized short videos, cinematic concept clips, lightweight motion storytelling

Sample Prompts

  1. A cinematic push-in shot as the subject slowly turns toward the camera, soft natural motion, subtle background movement, realistic lighting, polished commercial style
  2. A premium product on a clean tabletop, slow orbit camera movement, soft reflections, minimal background, elegant advertising look
  3. A character standing in a neon-lit street, slight hair and clothing motion, gentle forward camera move, atmospheric futuristic mood

💰 Pricing

ResolutionPrice per SecondInput Image Fee5s Example
480p$0.08$0.01 / image$0.41
720p$0.14$0.01 / image$0.71

Billing Rules

RuleDetails
Pricing modelLinear pricing based on output duration
Duration billingRounded up to the next whole second
Minimum billed Duration1 second
Maximum billed Duration15 seconds
Additional chargeEach request includes a fixed $0.01 input image fee

💡 Best Use Cases

  • Social media content: Animate still visuals into short promotional clips for posts, reels, and lightweight campaigns.
  • E-commerce and product marketing: Turn product images into motion-driven demos for ads, landing pages, and creative testing.
  • Concept visualization: Explore camera direction, motion ideas, and scene atmosphere from a single static frame.
  • Creative prototyping: Quickly test Prompt-based motion concepts for characters, environments, and branded storytelling.

🔗 Related Models

  • xAI text-to-image workflows: Useful when you want to generate the source image before animating it.
  • Other image-to-video workflows: Worth comparing when you need different trade-offs in quality, speed, or motion control.

Related models

vidu/q3/image-to-video-spicy
Image to VideoVidu

vidu/q3/image-to-video-spicy

Vidu Q3 Image-to-Video Spicy generates unlimited high-quality videos from images with smooth animations and diverse motion, optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.750 / per run
Vidu Q3 Image To Video
Image to VideoVidu

Vidu Q3 Image To Video

Vidu Q3 Image-to-Video turns text prompts into high-quality videos with exceptional visual fidelity and diverse motion. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.750 / per run
MiniMax H3 Spicy Image to Video
Image to VideoMinimax

MiniMax H3 Spicy Image to Video

Animate a first-frame image with optional prompt and last-frame guidance. Supports 480p, 540p, 768p and 1080p video, with durations from 3 to 15 seconds.

from $0.200 / per run
MiniMax H3 Image to Video
Image to VideoMinimax

MiniMax H3 Image to Video

MiniMax H3 Image to Video animates a first-frame image into a coherent 2K video, with natural-language motion instructions and optional last-frame control for consistent motion, scene continuity, and cinematic video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.650 / per run
ltx-2.5/image-to-video
Image to VideoLtx 2.5

ltx-2.5/image-to-video

LTX 2.5 Image-to-Video animates a first-frame image into high-fidelity synchronized audio-video content, with optional last-frame guidance and 720P / 1080P / 2K / 4K output for cinematic videos, social content, ads, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.500 / per run
ai/ltx-2.3-spicy/image-to-video-lora
Image to VideoLightricks

ai/ltx-2.3-spicy/image-to-video-lora

LTX 2.3 Spicy LoRA Image to Video API turns a reference image and prompt into expressive AI videos using selectable LoRA presets, optional LoRA strength overrides, duration, and resolution controls. Run fast REST inference on MaaS with no cold starts and affordable pricing.

from $0.150 / per run
ai/ltx-2.3-spicy/image-to-video
Image to VideoLightricks

ai/ltx-2.3-spicy/image-to-video

LTX 2.3 Spicy Image-to-Video API generates expressive AI videos from a reference image and prompt. Choose style presets, duration, and resolution with fast REST inference, no cold starts, and affordable pricing on MaaS.

from $0.100 / per run
Kling Omni Video O3 Image-To-Video
Image to VideoKling

Kling Omni Video O3 Image-To-Video

Kling Omni Video O3 Image-to-Video transforms static images into dynamic cinematic videos using MVL (Multi-modal Visual Language) technology. Maintains subject consistency while adding natural motion, physics simulation, and seamless scene dynamics. Supports audio generation. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.

from $0.420 / per run