Alibaba AI Models
Browse all 12 Alibaba AI model APIs on NamiFusion — image, video and face-swap models you can try in the playground and integrate through one unified REST API. Pay as you go, no arbitrary content filters, no forced watermark.
12 models
Wan 3.0 Omni Reference
Generates videos up to 30 seconds with native audio from multimodal references — images, video, audio, documents, and web pages. Reproduces characters, props, and spatial relationships from your references at pixel-level fidelity, with a first/last-frame mode for precise control.
Wan 3.0 Prime
Supports multimodal references including multiple images, video, and audio to generate videos with rich motion and synchronized audio-visual effects. Accurately preserves the appearance, motion characteristics, and visual style of referenced subjects, while supporting flexible combinations of references for versatile content creation and shot control.
alibaba/wan-2.7/text-to-image-pro
WAN 2.7 Text-to-Image Pro generates high-quality images up to 4K from text prompts with thinking mode for enhanced image quality. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
alibaba/wan-2.7/text-to-image
WAN 2.7 Text-to-Image generates high-quality images from text prompts with thinking mode for enhanced image quality. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
alibaba/wan-2.7/image-edit-pro
WAN 2.7 Image Edit Pro performs prompt-driven image editing with multi-image reference support and up to 2K output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
alibaba/wan-2.7/image-edit
WAN 2.7 Image Edit performs prompt-driven image editing with support for multiple-image references. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Alibaba WAN 2.6
Alibaba WAN 2.6 Reference-to-Video seamlessly transforms character, prop, or scene references—supporting both single and multi-view inputs—into high-quality video sequences. It excels at preserving identity, style, and layout while delivering fluid, coherent motion. Experience peak performance via our production-ready REST API, featuring zero cold starts and cost-effective pricing.
Alibaba WAN 2.6
Alibaba WAN 2.6 converts text or images into videos (720p/1080p) with synced audio, faster and more affordable than Google Veo3. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Alibaba WAN 2.5
Alibaba WAN 2.5 delivers high-fidelity text and image-to-video generation from 480p up to 1080p, featuring seamlessly synced audio. Faster and more cost-efficient than Google Veo 3, it offers a ready-to-use REST inference API with peak performance and zero cold starts. Achieve professional-grade video production with industry-leading speed and pricing.
Alibaba WAN 2.5
Alibaba WAN 2.5 redefines AI video generation by transforming text or images into high-quality videos (480p/720p/1080p) with natively synced audio. Engineered to be faster and more cost-effective than Google Veo 3, it offers a ready-to-use REST API with industry-leading performance and zero cold starts. It is the ultimate solution for creators seeking professional-grade output at a fraction of the cost.
Qwen Image 3.0 Text to Image
Qwen Image 3.0 is a professional-grade text-to-image model with superior quality and advanced prompt understanding. Up to 2K. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Qwen Image 3.0 Edit | Fast Image Editing
Qwen Image 3.0 Edit is a professional-grade image editing model with superior quality and advanced instruction understanding. Up to 2K. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.