How to Use the ai/heartmula/transcribe-lyrics API: speech to text integration guide (2026)

2 min · Updated 2026-07-11 · By the NamiFusion team

Learn how to call the ai/heartmula/transcribe-lyrics API on NamiFusion: parameters, pricing (from ≈$0.050) and runnable copy-paste code — try it in the playground first.

ai/heartmula/transcribe-lyrics is a speech to text model by Heartmula: HeartMuLa Transcribe extracts lyrics from audio files using advanced AI. Supports multilingual transcription. Ready-to-use REST inference API with best performance, no coldstarts, and affordable pricing. On NamiFusion you can try it in the playground (exact quote shown before you submit), then call it through one unified REST API — about $0.050 per typical run.

TL;DR

· Task type: speech to text, provided by Heartmula

· Pricing: about $0.050 per typical run ($1 = 100 credits), pay-as-you-go

· Getting started: playground (exact quote before submit) → create an API key → one POST request

· Tunable parameters: 1 (see the table below)

· No platform watermark; commercial use allowed (per the Terms of Service)

What can you build with ai/heartmula/transcribe-lyrics?

Voiceovers and narration: generate multi-language voices for short videos, courses and ads.

Audio content production: turn articles, news and fiction into publishable audio at scale.

Product voice interfaces: consistent branded voices for apps and devices.

How do you call this API?

  1. Sign up on NamiFusion and create a key on the API Keys page (signup bonus credits can only be used on selected models).
  2. Submit a task with the cURL example below, passing your key as a Bearer token.
  3. Poll the task status with the returned task_uuid and read the output URLs when completed.
Submit a task
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/heartmula/transcribe-lyrics" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "input": {
    "audio": "https://d5v2vcqcwe9y5.cloudfront.net/algorithm/video_translate/260327/default/qtqfcgymg08o.wav"
  }
}'

# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
  -H "Authorization: Bearer YOUR_API_KEY"

What parameters does it take?

audio (Audio): required

How much does it cost?

Pay-per-use in credits. About $0.050 per typical run ($1 = 100 credits); the exact price varies with parameters like resolution or duration, and the playground shows the exact quote before you submit.

ai/heartmula/transcribe-lyrics

Fast, multilingual lyric transcription with no cold starts and affordable per-run pricing.

ai/heartmula/transcribe-lyrics is an AI model designed to extract lyrics from audio files with a focus on speed, reliability, and cost efficiency. It supports multilingual transcription and is well suited for production workflows that require consistent performance and ready-to-use lyric extraction.

🚀 Key Features

  • Lyric-focused transcription: Built specifically for extracting sung lyrics from audio, making it a strong fit for music-centric workflows.
  • Multilingual support: Handles multilingual audio content for broader catalog coverage and international use cases.
  • No cold start: Delivers more consistent responsiveness for real-time and high-volume production environments.
  • Simple audio input: Accepts an audio URL, making it easy to plug into existing media pipelines.
  • Affordable pricing: Low per-run cost makes it practical for batch processing and large-scale lyric indexing.

🛠️ Technical Specifications

ItemDetails
Model architectureAI lyric transcription model
Task typeAudio-to-text(lyric extraction)
InputAudio URL
Input formatString(audio)
OutputTranscribed lyrics text
Language supportMultilingual
LatencyOptimized for online inference with no cold starts
Deployment profileProduction-ready performance

Sample Prompts

Note: This model primarily operates on audio input rather than Prompt-based control. The examples below illustrate common task intents.

  • Extract the full lyrics from this song audio.
  • Transcribe the lyrics from this multilingual music track.
  • Identify the sung content in this audio file and return readable lyrics text.

💰 Pricing

ItemPrice
Base Price$0.05

FAQ

Can I try it for free first?+

Signup bonus credits can only be used on selected models, and eligibility may change. The playground shows the exact quote before submission.

Can I use the output commercially? Is there a watermark?+

Yes, outputs are yours to use commercially, and NamiFusion adds no platform watermark — subject to the Terms of Service and applicable law.

What is the content policy?+

Pornographic content, sexualized content involving a minor, unauthorized use of real people, deceptive impersonation, fraud and false endorsements are prohibited. Lawful, consensual, non-explicit adult themes may be permitted; see the Terms of Service.

How long does a ai/heartmula/transcribe-lyrics task take?+

It depends on parameters — images usually take seconds, videos tens of seconds to minutes. The API returns a task_uuid for polling.

What are the resolution and duration limits?+

See the parameter table above — each model lists its available resolution and duration options there and in the playground.

How to Use the ai/heartmula/transcribe-lyrics API: speech to text integration guide (2026) | NamiFusion Blog