VT Lip Sync
namifusion/vt-lipsync
namifusion/vt-lipsync. Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.
Mô hình liên quan
Đầy đủ tham số, lược đồ đầu ra và giá hiện tại
| Chợ API mô hình AI | Một nền tảng cho mọi thứ | Phổ biến nhất | Điều khoản | Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất. |
|---|---|---|---|---|
| audio_url *Translated Audio | audio_upload | — | audio/mpeg,audio/wav,audio/ogg,audio/webm,audio/aac,audio/flac,audio/* | |
| subs_list *Translated Subtitle Segments | array<object> | [] | — | |
| ↳start_ms *Start Time (ms) | number | — | step 1 | |
| ↳end_ms *End Time (ms) | number | — | step 1 | |
| original_video_url *Original Video | video_upload | — | video/mp4,video/webm,video/quicktime,video/* | |
| original_subs_list *Original Subtitle Segments | array<object> | [] | — | |
| ↳start_ms *Start Time (ms) | number | — | step 1 | |
| ↳end_ms *End Time (ms) | number | — | step 1 |
Đầy đủ tham số, lược đồ đầu ra và giá hiện tại
| Tìm hiểu thêm | Một nền tảng cho mọi thứ | Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất. |
|---|---|---|
| lipsync_video_url | string |
API
Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.
cURL
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/namifusion/vt-lipsync" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"audio_url": "https://d5v2vcqcwe9y5.cloudfront.net/algorithm/video_translate/260327/default/qtqfcgymg08o.wav",
"subs_list": [
{
"end_ms": 4178,
"start_ms": 300
}
],
"original_video_url": "https://d5v2vcqcwe9y5.cloudfront.net/video_translate/260327/6964a3741d6212ca41d15d2d/3ftvqu8aewix.mp4",
"original_subs_list": [
{
"end_ms": 3780,
"start_ms": 300
}
]
}
}'
# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"Python
import time, requests
API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
# 1) Submit
resp = requests.post(
"https://www.namifusion.com/api/v1/marketplace/run/namifusion/vt-lipsync",
headers=HEADERS,
json={
"input": {
"audio_url": "https://d5v2vcqcwe9y5.cloudfront.net/algorithm/video_translate/260327/default/qtqfcgymg08o.wav",
"subs_list": [
{
"end_ms": 4178,
"start_ms": 300
}
],
"original_video_url": "https://d5v2vcqcwe9y5.cloudfront.net/video_translate/260327/6964a3741d6212ca41d15d2d/3ftvqu8aewix.mp4",
"original_subs_list": [
{
"end_ms": 3780,
"start_ms": 300
}
]
}
},
)
resp.raise_for_status() # 401/402/429/5xx stop here instead of polling a bad task
task = resp.json()
# 2) Poll until a terminal state (completed / failed / cancelled).
# This model is allowed up to 1800s server-side.
deadline = time.time() + 1860
while task.get("status") not in ("completed", "failed", "cancelled"):
if time.time() > deadline:
raise TimeoutError(f"still {task.get('status')} — keep the task_uuid and poll later")
time.sleep(3)
poll = requests.get(f"https://www.namifusion.com/api/v1/marketplace/run/tasks/{task['task_uuid']}", headers=HEADERS)
poll.raise_for_status()
task = poll.json()
print(task["status"], task.get("output"))JavaScript
const API_KEY = "YOUR_API_KEY";
const HEADERS = { Authorization: `Bearer ${API_KEY}` };
// 1) Submit
const resp = await fetch("https://www.namifusion.com/api/v1/marketplace/run/namifusion/vt-lipsync", {
method: "POST",
headers: { ...HEADERS, "Content-Type": "application/json" },
body: JSON.stringify({
"input": {
"audio_url": "https://d5v2vcqcwe9y5.cloudfront.net/algorithm/video_translate/260327/default/qtqfcgymg08o.wav",
"subs_list": [
{
"end_ms": 4178,
"start_ms": 300
}
],
"original_video_url": "https://d5v2vcqcwe9y5.cloudfront.net/video_translate/260327/6964a3741d6212ca41d15d2d/3ftvqu8aewix.mp4",
"original_subs_list": [
{
"end_ms": 3780,
"start_ms": 300
}
]
}
}),
});
if (!resp.ok) throw new Error(`submit failed: ${resp.status} ${await resp.text()}`);
let task = await resp.json();
// 2) Poll until a terminal state (completed / failed / cancelled).
// This model is allowed up to 1800s server-side.
const deadline = Date.now() + 1860 * 1000;
while (!["completed", "failed", "cancelled"].includes(task.status)) {
if (Date.now() > deadline) throw new Error(`still ${task.status} — keep the task_uuid and poll later`);
await new Promise((r) => setTimeout(r, 3000));
const poll = await fetch(`https://www.namifusion.com/api/v1/marketplace/run/tasks/${task.task_uuid}`, { headers: HEADERS });
if (!poll.ok) throw new Error(`poll failed: ${poll.status}`);
task = await poll.json();
}
console.log(task.status, task.output);Tài liệu API
NamiFusion VT LipSync
Lip-sync for translated videos: input a translated video, translated audio, and sentence-level timelines for both sides, then output a lip-synced result video.
NamiFusion VT LipSync is a lip-sync service designed for video translation scenarios. It is suitable for cases where the duration and rhythm of a video change after translation. By providing the translated video, translated audio, and two sets of sentence-level timing segments for the original and translated content, it generates a lip-synced video aligned with the target audio.
Key Features
- Built for video translation workflows: Specifically handles lip-sync drift caused by changes in audio and video length after translation.
- Dual timeline input: Accepts both translated sentence-level timelines and original sentence-level timelines to establish sentence-to-sentence alignment.
- Async Tasks: Returns a
task_uuidafter submission, and you can retrieve results through polling or Webhook. - Result video output: Returns the lip-synced video URL when the task is completed.
Technical Specifications
| Parameter | Details |
|---|---|
| Model ID | namifusion/vt-lipsync |
| Request Method | Async POST (submit task + poll/Webhook for results) |
| Input | Translated video URL + translated audio URL + dual sentence-level timelines |
| Output | Lip-synced video URL |
| Processing Time | Typically tens of seconds to several minutes, depending on video duration and number of sentence segments |
Quick Start
API Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/v1/marketplace/run/namifusion/vt-lipsync | POST | Submit a VT LipSync task |
/api/v1/marketplace/run/tasks/{task_uuid} | GET | Query task status and results |
Authentication
Include your API Key in the request header:
X-API-Key: sk-your-api-key
Usage Example
Step 1: Submit a Task
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/namifusion/vt-lipsync" \
-H "X-API-Key: sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"input": {
"video_url": "https://example.com/translated_video.mp4",
"audio_url": "https://example.com/translated_audio.wav",
"subs_list": [
{ "start_ms": 0, "end_ms": 8090 },
{ "start_ms": 8090, "end_ms": 10269 }
],
"original_video_url": "https://example.com/original_video.mp4",
"original_subs_list": [
{ "start_ms": 0, "end_ms": 7840 },
{ "start_ms": 7840, "end_ms": 9820 }
],
}
}'
Response Example:
{
"task_uuid": "69b931ac-d2db-d096-fc0a-bed1a2c3d4e5",
"status": "pending",
"cost_credits": 0
}
Step 2: Poll Task Status
curl -X GET "https://www.namifusion.com/api/v1/marketplace/run/tasks/69b931ac-d2db-d096-fc0a-bed1a2c3d4e5" \
-H "X-API-Key: sk-your-api-key"
Processing:
{
"task_uuid": "69b931ac-d2db-d096-fc0a-bed1a2c3d4e5",
"status": "processing"
}
Completed:
{
"task_uuid": "69af63b1-285a-4ae0-fe92-6cc5a1b2c3d4",
"model_id": "namifusion/vt-lipsync",
"status": "completed",
"output": {
"lipsync_video_url": "https://d5v2vcqcwe9y5.cloudfront.net/algorithm/lipsync/260310/default/tx6hbykt7ntp.mp4",
"lipsync_from": 1
},
"cost_credits": 0,
"created_at": "2026-03-08T00:20:01Z",
"completed_at": "2026-03-08T00:21:31Z"
}
Python Example
import requests
import time
API_KEY = "sk-your-api-key"
BASE_URL = "https://www.namifusion.com/api/v1/marketplace/run"
HEADERS = {
"X-API-Key": API_KEY,
"Content-Type": "application/json",
}
payload = {
"input": {
"video_url": "https://example.com/translated_video.mp4",
"audio_url": "https://example.com/translated_audio.wav",
"subs_list": [
{"start_ms": 0, "end_ms": 8090},
{"start_ms": 8090, "end_ms": 10269},
],
"original_video_url": "https://example.com/original_video.mp4",
"original_subs_list": [
{"start_ms": 0, "end_ms": 7840},
{"start_ms": 7840, "end_ms": 9820},
],
}
}
task_resp = requests.post(
f"{BASE_URL}/namifusion/vt-lipsync",
headers=HEADERS,
json=payload,
).json()
task_uuid = task_resp["task_uuid"]
print(f"Task submitted: {task_uuid}")
while True:
status_resp = requests.get(
f"{BASE_URL}/tasks/{task_uuid}",
headers=HEADERS,
).json()
status = status_resp["status"]
print(f"Status: {status}")
if status == "completed":
result = status_resp["output"]
print("Result video:", result["lipsync_video_url"])
print("Source type:", result.get("lipsync_from"))
break
if status == "failed":
print("Failed:", status_resp.get("error_message", "Unknown error"))
break
time.sleep(5)
Detailed Parameters and Return Values
Request Parameters
The request body is JSON, and all parameters are placed inside the input object:
{
"input": {
"video_url": "https://...",
"audio_url": "https://...",
"subs_list": [
{ "start_ms": 6140, "end_ms": 8090 }
],
"original_video_url": "https://...",
"original_subs_list": [
{ "start_ms": 6140, "end_ms": 7840 }
]
}
}
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
video_url | string | Yes | - | URL of the translated video. This video has already been retimed or duration-adjusted for the target language. |
audio_url | string | Yes | - | URL of the translated audio. The service performs lip-sync based on this audio. |
subs_list | array | Yes | - | ASR sentence segmentation result for the translated audio. Each item represents the start and end time of one sentence. |
original_video_url | string | Yes | - | URL of the original video uploaded by the user. Used as the source reference for original mouth movements. |
original_subs_list | array | Yes | - | ASR sentence segmentation result for the original video audio. Each item represents the start and end time of one sentence. |
Sentence Timeline Object
Each item in subs_list and original_subs_list uses the same structure:
| Field | Type | Required | Description |
|---|---|---|---|
start_ms | integer | Yes | Start time of the current sentence, in milliseconds. |
end_ms | integer | Yes | End time of the current sentence, in milliseconds. |
Parameter Constraints and Correspondence
| Item | Description |
|---|---|
subs_list and original_subs_list | These are the ASR sentence segmentation results for the translated audio and original audio respectively. Their lengths do not need to match. |
| Time unit | All time values use milliseconds. |
| Time order | Each sentence object should satisfy start_ms < end_ms. |
| URL accessibility | All URLs should be publicly accessible so the server can download them directly. |
| Audio-video correspondence | video_url and audio_url must belong to the same translated content; original_video_url and original_subs_list must belong to the same original content. |
Task Status
After submitting the task, query its status by polling. A 5-second polling interval is recommended:
| Status | Description |
|---|---|
pending | The task has been created and is waiting to be processed. |
processing | The task is being processed. |
completed | The task is completed, and the result is available in output.lipsync_video_url. |
failed | The task failed. Check error_message for the reason. |
Return Value Structure
Task Submission Response
| Field | Type | Description |
|---|---|---|
task_uuid | string | Unique task identifier used for later status queries. |
status | string | Initial task status, usually pending. |
cost_credits | number | Credits consumed by this task. |
Task Completion Response
| Field | Type | Description |
|---|---|---|
task_uuid | string | Unique task identifier. |
model_id | string | Model ID (namifusion/vt-lipsync). |
status | string | Task status. |
output.lipsync_video_url | string | URL of the lip-synced video. |
output.lipsync_from | integer | Result source marker. The current example response returns 1. |
cost_credits | number | Credits consumed. |
created_at | string | Task creation time (ISO 8601). |
completed_at | string | Task completion time (ISO 8601). |
error_message | string | Error message, returned only when failed. |
Internal Task Field Mapping
The following table shows the relationship between the public API fields and the underlying task document fields:
| Public Request Field | Underlying Task Document Field |
|---|---|
input.video_url | extra.video_url |
input.audio_url | extra.audio_url |
input.subs_list | extra.subs_list |
input.original_video_url | extra.original_video_url |
input.original_subs_list | extra.original_subs_list |
output.lipsync_video_url | data.lipsync_video_url |
output.lipsync_from | data.lipsync_from |
How Parameters Affect the Output
| Parameter | Effect on Output |
|---|---|
video_url | Determines the base visuals and translated-video pacing used for lip-sync alignment. |
audio_url | Determines the target speech content that the result video must align with. |
subs_list | ASR sentence segmentation result for the translated audio. It defines the sentence boundaries in the translated content and directly affects sentence-level lip-sync realignment. |
original_video_url | Provides the original mouth movement reference to preserve the speaking characteristics of the source video. |
original_subs_list | ASR sentence segmentation result for the original audio. It is used to map sentence-level timing from the original video to the translated version. |
Error Response
When a task fails, the polling response contains an error message:
{
"task_uuid": "69b931ac-d2db-d096-fc0a-bed1a2c3d4e5",
"status": "failed",
"error_message": "Invalid subtitle alignment"
}
Common errors:
| Error Code | Description |
|---|---|
| 400 | Invalid parameters, such as missing required fields or invalid timeline format. |
| 401 | Authentication failed because the API Key is missing or invalid. |
| 408 | Timeout while downloading video or audio. |
| 422 | Input media is inaccessible, or the sentence timelines do not match the content correctly. |
| 500 | Internal service error. |
Notes
- Provide both the translated video and translated audio: The service depends on both inputs to perform lip-sync correctly.
- Sentence timelines must come from ASR results of their respective audio tracks:
subs_listis the ASR sentence segmentation result for the translated audio, andoriginal_subs_listis the ASR sentence segmentation result for the original audio. Their lengths do not need to match. - Translated videos may have different durations: This service is designed for scenarios where the translated video length differs from the original video.
- Stable sentence segmentation is recommended: Segmentation that is too coarse or too fragmented may reduce sync quality.
- Result field: The final output video URL is located at
output.lipsync_video_url.
Mô hình liên quan
Kling Omni Video O3
kwaivgi/kling-video-o3/video-edit. Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.
Kling Omni Video O3
kwaivgi/kling-video-o3/reference-to-video. Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.
Kling 3.0 Standard
kwaivgi/kling-v3.0/motion-control. Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.
seedance-2-0 reference-to-video
doubao/seedance-2-0/reference-to-video. Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.
Alibaba WAN 2.6
alibaba/wan-2.6/reference-to-video. Duyệt API mô hình ảnh, video và đổi khuôn mặt. Thử mô hình rồi tích hợp qua một API REST thống nhất.