How to Use the MiniMax H3 Text to Video API: text to video integration guide (2026)

2 min · Updated 2026-07-11 · By the NamiFusion team

Learn how to call the MiniMax H3 Text to Video API on NamiFusion: parameters, pricing (from ≈$0.650) and runnable copy-paste code — try it in the playground first.

MiniMax H3 Text to Video is a text to video model by Minimax: MiniMax H3 Text to Video generates coherent 2K videos from text prompts, with flexible 4-15 second duration and six selectable aspect ratios for cinematic scenes, creative videos, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. On NamiFusion you can try it in the playground (exact quote shown before you submit), then call it through one unified REST API — a typical run costs about $0.650.

TL;DR

· Task type: text to video, provided by Minimax

· Pricing: about $0.650 per typical run ($1 = 100 credits), pay-as-you-go

· Getting started: playground (exact quote before submit) → create an API key → one POST request

· Tunable parameters: 4 (see the table below)

· No platform watermark; commercial use allowed (per the Terms of Service)

What can you build with MiniMax H3 Text to Video?

Short-form video and ad clips: go from a prompt or a single image to publishable motion content.

Product demos and concept previews: validate storyboards and visual direction cheaply before any shoot.

A social content factory: batch vertical clips and ride trending topics fast.

Remixing existing assets: image-to-video brings static work to life and extends its lifespan.

How do you call this API?

1. Sign up on NamiFusion and create a key on the API Keys page (signup bonus credits can only be used on selected models).

2. Submit a task with the cURL example below, passing your key as a Bearer token.

3. Poll the task status with the returned task_uuid and read the output URLs when completed.

Submit a task
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/minimax/h3/text-to-video" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "input": {
    "prompt": "A cinematic portrait, dramatic rim light, shallow depth of field",
    "aspect_ratio": "16:9",
    "resolution": "2k",
    "duration": 5
  }
}'

# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"

What parameters does it take?

prompt (Prompt): required

aspect_ratio (Aspect Ratio): optional, default "16:9"

resolution (Resolution): optional, default "2k"

duration (Duration): optional, default 5

How do you write prompts for MiniMax H3 Text to Video?

1. Subject first: lead with the actor and action, then add environment, lighting and style ("an astronaut sprinting through a neon street in the rain, cinematic, shallow depth of field").

2. Concrete nouns beat stacked adjectives: "85mm portrait lens, golden-hour backlight" works far better than "very beautiful, ultra high quality".

3. Change one variable per iteration: pin the seed to reproduce a result, tweak a single phrase at a time so you know what actually moved the output.

4. State negatives positively: instead of "no blur", ask for "sharp focus, crisp details".

How much does it cost?

Pay-per-use in credits. A typical run costs about $0.650 ($1 = 100 credits); the exact price varies with parameters like resolution or duration, and the playground shows the exact quote before you submit.

MiniMax H3 Text to Video

Affordable 2K text-to-video generation with flexible durations, multiple aspect ratios, and no cold start.

MiniMax H3 Text to Video is a production-ready model that generates coherent 2K videos from natural-language prompts. It is designed for creative and commercial workflows, offering flexible 4–15 second outputs, broad aspect ratio support, and reliable performance for rapid iteration.

🚀 Key Features

  • 2K video output: Generate high-resolution videos with a fixed 2K output tier for strong visual clarity.
  • Text-to-video generation: Turn natural-language descriptions into videos by specifying scene, action, camera movement, and style.
  • Flexible duration control: Choose output lengths from 4 to 15 seconds to match testing, prototyping, or final content needs.
  • Multiple aspect ratios: Supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 for cinematic, social, and mobile-first formats.
  • No cold start: Well suited for production workflows that require consistent responsiveness.
  • Transparent pricing: Per-second pricing makes cost planning straightforward across different video lengths.

🛠️ Technical Specifications

ItemDetails
Model NameMiniMax H3 Text to Video
Model IDminimax/h3/text-to-video
Model ArchitectureText-to-video generation model
Input FormatText prompt
Output FormatVideo
Resolution2K
Duration4–15 seconds
Aspect Ratio21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Default Aspect Ratio16:9
Default Resolution2K
Default Duration5 seconds
Prompt Length1–7000 characters
LatencyNo cold start; actual inference time varies by prompt complexity and duration
Typical WorkflowsCinematic scenes, creative video generation, marketing content, production previsualization

Sample Prompts

  • A cinematic ocean wave at sunrise, highly detailed, golden light breaking through clouds, slow push-in camera movement, epic atmosphere.
  • A fashion model walking through a futuristic city street at night, neon reflections on wet pavement, smooth tracking shot, premium commercial style.
  • A cozy cabin in a misty forest at dawn, sunlight filtering through trees, slow zoom from wide shot to medium shot, realistic and tranquil mood.

💰 Pricing

ModePrice
Base Price$0.13 / second
5-second video$0.65
10-second video$1.30
15-second video$1.95

FAQ

Can I try it for free first?+

Signup bonus credits can only be used on selected models, and eligibility may change. The playground shows the exact quote before submission.

Can I use the output commercially? Is there a watermark?+

Yes, outputs are yours to use commercially, and NamiFusion adds no platform watermark — subject to the Terms of Service and applicable law.

What is the content policy?+

NamiFusion is an unfiltered engine: no arbitrary restrictions and NSFW-friendly, but illegal content and non-consensual use of real people are prohibited — see the Terms of Service.

How long does a MiniMax H3 Text to Video task take?+

It depends on parameters — images usually take seconds, videos tens of seconds to minutes. The API returns a task_uuid for polling.

What are the resolution and duration limits?+

See the parameter table above — each model lists its available resolution and duration options there and in the playground.

How to Use the MiniMax H3 Text to Video API: text to video integration guide (2026) | NamiFusion Blog