MiniMax H3 Image to Video

minimax/h3/image-to-video

MiniMax H3 Image to Video animates a first-frame image into a coherent 2K video, with natural-language motion instructions and optional last-frame control for consistent motion, scene continuity, and cinematic video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Examples

MiniMax H3 Image to Video Examples 1

Parameters

NameTypeDefaultConstraintsDescription
first_frame *Imageimage_upload—image/* · 0–1 itemsFirst-frame image URL or Base64 data. Supported image dimensions are 256-5760 pixels per side, with an aspect ratio between 0.4 and 2.5.
prompt *Prompttextarea—≤ 7000 charsText description of the desired motion and scene.
last_frameLast Imageimage_upload—image/* · 0–1 itemsOptional last-frame image URL or Base64 data.
resolutionResolutionselect2k2kOutput video resolution.
durationDurationslider54 ~ 15 · step 1Output video duration in seconds.

Output fields

FieldTypeDescription
videosarray<string>Generated video URLs

API

Call this model through one unified REST API. Get a key on the API Keys page.

cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/minimax/h3/image-to-video" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "input": {
    "first_frame": [
      "https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
    ],
    "prompt": "Animate the scene with natural subject movement, gentle environmental motion, a stable camera path, and cinematic lighting.",
    "resolution": "2k",
    "duration": 5
  }
}'

# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"
Python
import time, requests

API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}

# 1) Submit
resp = requests.post(
    "https://www.namifusion.com/api/v1/marketplace/run/minimax/h3/image-to-video",
    headers=HEADERS,
    json={
        "input": {
            "first_frame": [
                "https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
            ],
            "prompt": "Animate the scene with natural subject movement, gentle environmental motion, a stable camera path, and cinematic lighting.",
            "resolution": "2k",
            "duration": 5
        }
    },
)
resp.raise_for_status()  # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()

# 2) Poll until a terminal state (completed / failed / cancelled).
#    This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
    if time.time() > deadline:
        raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
    time.sleep(3)
    poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
    poll.raise_for_status()
    task = poll.json()

print(task["status"], task.get("output"))
JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };

// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/minimax/h3/image-to-video", {
  method: "POST",
  headers: { ...HEADERS, "Content-Type": "application/json" },
  body: JSON.stringify({
    "input": {
      "first_frame": [
        "https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
      ],
      "prompt": "Animate the scene with natural subject movement, gentle environmental motion, a stable camera path, and cinematic lighting.",
      "resolution": "2k",
      "duration": 5
    }
  }),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();

// 2) Poll until a terminal state (completed / failed / cancelled).
//    This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
  if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
  await new Promise((r) => setTimeout(r, 3000));
  const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
  if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
  task = await poll.json();
}

console.log(task.status, task.output);

Documentation

MiniMax H3 Image to Video

Turn a single first-frame image into a coherent 2K video with prompt-driven motion control.

MiniMax H3 Image to Video is an image-to-video model that transforms a starting image into a high-resolution 2K video using natural-language motion instructions. With optional last-frame guidance, it is well suited for workflows that require stronger scene continuity, controlled progression, and cinematic short-form video generation. It also stands out for no cold start behavior and cost-efficient pricing.

🚀 Key Features

  • First-frame guided generation: Uses the input image to anchor subject identity, composition, style, and opening visual state.
  • Prompt-based motion control: Define movement, camera behavior, lighting, mood, and scene progression with natural language.
  • Optional last-frame guidance: Add a final reference image to better control the ending frame and improve visual continuity.
  • Native 2K output: Generates high-resolution 2K video suitable for premium creative and commercial content.
  • Flexible duration options: Supports video lengths from 4 to 15 seconds for both rapid iteration and more developed scenes.
  • No cold start: Designed for responsive production usage with more predictable turnaround time.

🛠️ Technical Specifications

ItemDetails
Model NameMiniMax H3 Image to Video
Model IDminimax/h3/image-to-video
Task Typeimage-to-video
Model ArchitectureFirst-frame conditioned video generation model with optional last-frame control
Input FormatImage(URL or Base64) + Prompt; optional last-frame image(URL or Base64)
Output FormatVideo
Input Image Requirements256–5760 pixels per side; supported aspect ratio from 0.4 to 2.5
Prompt Length1–7000 characters
Resolution2K
Duration4–15 seconds
FpsNot publicly specified
Last-frame ControlSupported(optional)
LatencyNo cold start; actual inference time varies by workload and duration

Sample Prompts

  1. Slow cinematic push-in on a person standing in a rainy neon-lit street, subtle head movement, reflections on wet pavement, realistic lighting, moody atmosphere.
  2. Animate the product with a slow rotating motion on a clean studio background, slight camera orbit, premium commercial lighting, emphasize metallic texture and highlights.
  3. A white cat jumps down from a windowsill and walks toward the camera, warm sunlight filling the room, natural motion, soft and cozy mood.

💰 Pricing

OptionPrice
Base Price$0.13 / second
5s video$0.65
10s video$1.30
15s video$1.95

💡 Best Use Cases

  • E-commerce and product marketing: Turn product images into dynamic showcase videos for ads, landing pages, and catalog content.
  • Social media content creation: Convert static visuals into short-form motion assets with stronger engagement potential.
  • Previsualization and concept development: Animate concept art or storyboard frames into short cinematic previews.
  • Character and scene animation: Bring illustrations, portraits, or environment art to life while preserving the original visual direction.

🔗 Related Models

  • MiniMax H3 Text-to-Video: Generates high-resolution video directly from Prompt input.
  • MiniMax H3 Reference-to-Video: Generates high-resolution video from reference media plus Prompt guidance.

Related models

xAI Grok Imagine Video v1.5 Image to Video
Image to VideoX Ai

xAI Grok Imagine Video v1.5 Image to Video

Animate one input image with a text prompt into a 1-15 second video at 480p or 720p.

from $0.850 / per run
vidu/q3/image-to-video-spicy
Image to VideoVidu

vidu/q3/image-to-video-spicy

Vidu Q3 Image-to-Video Spicy generates unlimited high-quality videos from images with smooth animations and diverse motion, optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.750 / per run
Vidu Q3 Image To Video
Image to VideoVidu

Vidu Q3 Image To Video

Vidu Q3 Image-to-Video turns text prompts into high-quality videos with exceptional visual fidelity and diverse motion. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.750 / per run
MiniMax H3 Spicy Image to Video
Image to VideoMinimax

MiniMax H3 Spicy Image to Video

Animate a first-frame image with optional prompt and last-frame guidance. Supports 480p, 540p, 768p and 1080p video, with durations from 3 to 15 seconds.

from $0.200 / per run
ltx-2.5/image-to-video
Image to VideoLtx 2.5

ltx-2.5/image-to-video

LTX 2.5 Image-to-Video animates a first-frame image into high-fidelity synchronized audio-video content, with optional last-frame guidance and 720P / 1080P / 2K / 4K output for cinematic videos, social content, ads, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

from $0.500 / per run
ai/ltx-2.3-spicy/image-to-video-lora
Image to VideoLightricks

ai/ltx-2.3-spicy/image-to-video-lora

LTX 2.3 Spicy LoRA Image to Video API turns a reference image and prompt into expressive AI videos using selectable LoRA presets, optional LoRA strength overrides, duration, and resolution controls. Run fast REST inference on MaaS with no cold starts and affordable pricing.

from $0.150 / per run
ai/ltx-2.3-spicy/image-to-video
Image to VideoLightricks

ai/ltx-2.3-spicy/image-to-video

LTX 2.3 Spicy Image-to-Video API generates expressive AI videos from a reference image and prompt. Choose style presets, duration, and resolution with fast REST inference, no cold starts, and affordable pricing on MaaS.

from $0.100 / per run
Kling Omni Video O3 Image-To-Video
Image to VideoKling

Kling Omni Video O3 Image-To-Video

Kling Omni Video O3 Image-to-Video transforms static images into dynamic cinematic videos using MVL (Multi-modal Visual Language) technology. Maintains subject consistency while adding natural motion, physics simulation, and seamless scene dynamics. Supports audio generation. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.

from $0.420 / per run