How to Use the Qwen Image 3.0 Text to Image API: text to image integration guide (2026)
Learn how to call the Qwen Image 3.0 Text to Image API on NamiFusion: parameters, pricing (from ≈$0.030) and runnable copy-paste code — try it in the playground first.
Qwen Image 3.0 Text to Image is a text to image model by Alibaba: Qwen Image 3.0 is a professional-grade text-to-image model with superior quality and advanced prompt understanding. Up to 2K. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. On NamiFusion you can try it in the playground (exact quote shown before you submit), then call it through one unified REST API — a typical run costs about $0.030.
TL;DR
· Task type: text to image, provided by Alibaba
· Pricing: about $0.030 per typical run ($1 = 100 credits), pay-as-you-go
· Getting started: playground (exact quote before submit) → create an API key → one POST request
· Tunable parameters: 6 (see the table below)
· No platform watermark; commercial use allowed (per the Terms of Service)
What can you build with Qwen Image 3.0 Text to Image?
Social media visuals and covers: batch-generate on-theme images and schedule them straight away.
E-commerce and ad creatives: product scenes, banner backgrounds, and multi-variant visuals for A/B tests.
Concept art and ideation: characters, environments and style boards, iterated cheaply and fast.
Pipeline chaining: outputs feed directly into image-to-video or face swap as source material.
How do you call this API?
1. Sign up on NamiFusion and create a key on the API Keys page (signup bonus credits can only be used on selected models).
2. Submit a task with the cURL example below, passing your key as a Bearer token.
3. Poll the task status with the returned task_uuid and read the output URLs when completed.
# 1) Submit — returns { "task_uuid": "..." }
curl -X POST "https://www.namifusion.com/api/v1/marketplace/run/alibaba/qwen-image-3.0/text-to-image" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"prompt": "A cinematic portrait, dramatic rim light, shallow depth of field",
"n": 1,
"prompt_extend": true
}
}'
# 2) Poll until status is "completed", then read the output URLs
curl "https://www.namifusion.com/api/v1/marketplace/run/tasks/TASK_UUID" \
-H "Authorization: Bearer YOUR_API_KEY"What parameters does it take?
prompt (Prompt): required
size (Size): optional
n (Number of Images): optional, default 1
negative_prompt (Negative Prompt): optional
prompt_extend (Prompt Extend): optional, default true
seed (Seed): optional
How do you write prompts for Qwen Image 3.0 Text to Image?
1. Subject first: lead with the actor and action, then add environment, lighting and style ("an astronaut sprinting through a neon street in the rain, cinematic, shallow depth of field").
2. Concrete nouns beat stacked adjectives: "85mm portrait lens, golden-hour backlight" works far better than "very beautiful, ultra high quality".
3. Change one variable per iteration: pin the seed to reproduce a result, tweak a single phrase at a time so you know what actually moved the output.
4. State negatives positively: instead of "no blur", ask for "sharp focus, crisp details".
How much does it cost?
Pay-per-use in credits. A typical run costs about $0.030 ($1 = 100 credits); the exact price varies with parameters like resolution or duration, and the playground shows the exact quote before you submit.
Qwen Image 3.0 Text to Image
Professional-grade text-to-image generation with strong prompt understanding and reliable performance.
Qwen Image 3.0 Text to Image is a production-ready text-to-image model designed for high-quality visual generation from natural language prompts. It combines superior image quality, advanced prompt understanding, support for both Chinese and English, and up to 2K output, making it a strong choice for teams that need consistent results, no cold starts, and cost-efficient scaling.
🚀 Key Features
- Professional-grade image quality: Built for high-fidelity visual generation suitable for commercial and creative workflows.
- Advanced prompt understanding: Interprets detailed and nuanced prompts more accurately for better alignment with user intent.
- Up to 2K output: Generates images up to 2048×2048 for sharper and more presentation-ready results.
- Batch generation: Produces up to 6 images in a single request for faster ideation and selection.
- Smart resolution: Automatically selects the best resolution from your prompt when no size is specified.
- Bilingual prompt support: Accepts both Chinese and English prompts, enabling flexible multilingual creation.
- No cold starts: Optimized for responsive, ready-to-use inference with more predictable latency.
- Affordable pricing: Delivers a strong balance of quality, reliability, and cost for scalable usage.
🛠️ Technical Specifications
| Item | Details |
|---|---|
| Model architecture | Professional-grade text-to-image model |
| Task type | Text-to-image generation |
| Input | Text prompt |
| Output Format | Image |
| Prompt | Supports Chinese and English |
| Prompt rewriting | Automatic prompt optimization (enabled by default) |
| Resolution | 512×512 up to 2048×2048; auto-selected by the model when unspecified |
| Size format | Specified as width*height, such as 10241024 or 1280720 |
| Images per request | 1–6 |
| Seed | Optional; 0–2147483647 for reproducibility, random when unspecified |
| Quality positioning | Professional-grade, high-fidelity image generation |
| Latency | Optimized for responsive performance with no cold start |
Sample Prompts
- Chinese: 一张电影感十足的未来城市夜景,霓虹灯反射在雨后的街道上,超高细节,真实光影
- Chinese: 极简风产品摄影,一只白色陶瓷杯置于浅灰背景中央,柔和棚拍光,干净构图
- English: A cinematic portrait of a traveler standing in a desert at sunset, dramatic lighting, ultra-detailed, realistic style
💰 Pricing
| Item | Price |
|---|---|
| Generated images (any resolution) | $0.03 / image |
Billed per generated image: a request producing n images is charged $0.03 × n. Unspecified size (smart resolution) is billed at the same rate.
FAQ
Can I try it for free first?+
Signup bonus credits can only be used on selected models, and eligibility may change. The playground shows the exact quote before submission.
Can I use the output commercially? Is there a watermark?+
Yes, outputs are yours to use commercially, and NamiFusion adds no platform watermark — subject to the Terms of Service and applicable law.
What is the content policy?+
NamiFusion is an unfiltered engine: no arbitrary restrictions and NSFW-friendly, but illegal content and non-consensual use of real people are prohibited — see the Terms of Service.
How long does a Qwen Image 3.0 Text to Image task take?+
It depends on parameters — images usually take seconds, videos tens of seconds to minutes. The API returns a task_uuid for polling.
What are the resolution and duration limits?+
See the parameter table above — each model lists its available resolution and duration options there and in the playground.