kwaivgi/kling-v2-ai-avatar-pro
kwaivgi/kling-v2-ai-avatar-pro
kwaivgi/kling-v2-ai-avatar-pro. Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.
Mô hình liên quan
Đầy đủ tham số, lược đồ đầu ra và giá hiện tại
| Chợ API mô hình AI | Một nền tảng cho mọi thứ | Phổ biến nhất | Điều khoản | Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất. |
|---|---|---|---|---|
| image *Image | image_upload | — | image/* · 0–1 items | |
| promptPrompt | textarea | — | ≤ 5000 chars |
Đầy đủ tham số, lược đồ đầu ra và giá hiện tại
| Tìm hiểu thêm | Một nền tảng cho mọi thứ | Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất. |
|---|---|---|
| videos | array<string> |
API
Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.
cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/kwaivgi/kling-v2-ai-avatar-pro" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"prompt": "The person speaks naturally in sync with the supplied audio, with subtle facial expressions, gentle head movements, and a steady camera. Preserve the appearance of the reference portrait."
}
}'
# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"Python
import time, requests
API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
# 1) Submit
resp = requests.post(
"https://www.namifusion.com/api/v1/marketplace/run/kwaivgi/kling-v2-ai-avatar-pro",
headers=HEADERS,
json={
"input": {
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"prompt": "The person speaks naturally in sync with the supplied audio, with subtle facial expressions, gentle head movements, and a steady camera. Preserve the appearance of the reference portrait."
}
},
)
resp.raise_for_status() # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()
# 2) Poll until a terminal state (completed / failed / cancelled).
# This model is allowed up to 300s server-side.
deadline = time.time() + 360
while task.get("status") not in ("completed", "failed", "cancelled"):
if time.time() > deadline:
raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
time.sleep(3)
poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
poll.raise_for_status()
task = poll.json()
print(task["status"], task.get("output"))JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };
// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/kwaivgi/kling-v2-ai-avatar-pro", {
method: "POST",
headers: { ...HEADERS, "Content-Type": "application/json" },
body: JSON.stringify({
"input": {
"image": [
"https://assets-public.namifusion.com/marketplace/images/2026-06-12/638e1ee18ac7.jpeg"
],
"prompt": "The person speaks naturally in sync with the supplied audio, with subtle facial expressions, gentle head movements, and a steady camera. Preserve the appearance of the reference portrait."
}
}),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();
// 2) Poll until a terminal state (completed / failed / cancelled).
// This model is allowed up to 300s server-side.
const deadline = Date.now() + 360 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
await new Promise((r) => setTimeout(r, 3000));
const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
task = await poll.json();
}
console.log(task.status, task.output);Tài liệu API
kwaivgi/kling-v2-ai-avatar-pro
Create high-quality talking-head avatar videos from a single portrait and your own audio.
kwaivgi/kling-v2-ai-avatar-pro is an image-to-video model designed for AI avatar generation. It turns a single portrait image and an audio track into a clean, stable, lip-synced talking-head video with strong identity consistency. With no cold start, strong output quality, and affordable usage-based pricing, it is well suited for social content, profile videos, intros, and virtual presenter workflows.
🚀 Key Features
- Audio-driven lip sync: Uses uploaded audio directly rather than synthetic speech, preserving timing, pauses, and emotional delivery.
- Strong identity consistency: Maintains facial characteristics from the reference image while animating the face, eyes, and head naturally.
- Stable talking-head motion: Optimized for clean, on-camera avatar performance with controlled movement and reliable visual coherence.
- One-shot workflow: Requires only one Image and one Audio input, eliminating the need for video capture or motion recording.
- Prompt-guided styling: Supports optional Prompt input to influence mood, expression, lighting feel, or subtle camera presence.
- Social-ready vertical output: Ideal for short-form vertical content formats used across TikTok, Reels, Shorts, and Stories.
🛠️ Technical Specifications
| Item | Details |
|---|---|
| Model Name | kwaivgi/kling-v2-ai-avatar-pro |
| Model Type | Image-to-video(AI avatar / talking-head generation) |
| Core Function | Generates a lip-synced avatar Video from a single portrait Image and an Audio track |
| Inputs | Image, Audio, optional Prompt |
| Output Format | Video |
| Image Requirements | Clear portrait, preferably front-facing or 3/4 view, visible eyes, minimal occlusion |
| Audio Requirements | Clean mono or stereo speech audio with limited background noise |
| Duration | Automatically derived from the input audio length; minimum billed duration is 5 seconds |
| Resolution | HD vertical output(exact pixel dimensions not specified in source material) |
| Aspect Ratio | Vertical / portrait-oriented social format |
| Motion Driver | Audio-driven lip sync, timing, and facial performance |
| Style Control | Optional Prompt for expression, mood, lighting, and subtle motion cues |
| Latency | No fixed public value provided; platform highlights no cold start |
| Output Characteristics | Clean detail, stable motion, strong identity preservation |
Sample Prompts
soft studio lighting, subtle head movement, gentle smileconfident presenter in a tech promo, subtle head nodsfriendly customer service tone, warm expression, clean professional framing
💰 Pricing
| Audio Length(s) | Billed Seconds | Price(USD) |
|---|---|---|
| 0–5 | 5 | 0.56 |
| 10 | 10 | 1.12 |
| 20 | 20 | 2.24 |
| 30 | 30 | 3.36 |
| 60 | 60 | 6.72 |
Clips shorter than 5 seconds are still billed as 5 seconds.
💡 Best Use Cases
- Social media avatar content: Produce short-form talking-head clips for TikTok, Reels, Shorts, and similar channels.
- Brand intros and profile videos: Create polished avatar-based introductions, spokesperson clips, and profile content.
- Education and explainers: Build virtual presenter videos for tutorials, onboarding, and training materials.
- Marketing and customer engagement: Generate avatar-led promos, customer support explainers, and UGC-style ad creatives.
🔗 Related Models
- infinitetalk: Designed for lip-synced talking-head avatar generation from scripts or audio, suitable for virtual presenters and explainer content.
- Infinitetalk-Multi: Extends avatar generation to multi-speaker or multi-segment scenarios such as dialogues and batch content production.
- Omni-Human: Focuses on high-fidelity digital humans for virtual hosts, brand ambassadors, and training avatars.