Alibaba releases 8-step Qwen-Image-2.1-Turbo model
Alibaba’s Qwen team has launched Qwen-Image-2.1-Turbo, an 8-step image model that drastically cuts generation steps while retaining 2K resolution and multi-reference editing features.

Alibaba’s Qwen team has released Qwen-Image-2.1-Turbo, an accelerated open-weight checkpoint that reduces denoising requirements from 40 steps down to 8 steps. Built on a 7B parameter single-stream DiT visual generator with 32 layers and block-causal attention, the model is paired with a Qwen3-VL 8B text encoder and a 64-channel RGBA autoencoder featuring 16x spatial compression. The update preserves full 2K output capabilities, supporting square resolutions up to 2048x2048 and 16:9 aspect ratios at 2752x1536, alongside native transparent image generation.
The architecture relies on prefix KV caching to store context from input images and text during the initial step, reusing it across the remaining seven steps to reduce computational overhead. It uses Flow Matching with Euler discrete scheduling, dynamic shifting, and operates at CFG=1 by default. Practitioners can perform text-to-image generation and complex editing tasks, including multi-reference composition supporting up to 10 reference images and local edits via masks or painted annotations. Running the model locally requires Diffusers from source with PR #14950 and transformers version 5.17.0 or higher in BF16 on CUDA GPUs. While specific VRAM minimums for Turbo were not disclosed, estimates from Unsloth for the base model range from 11 GB using GGUF to 24 GB for INT8 or FP8 precision.
While Turbo lacks its own evaluation benchmark, the underlying base Qwen-Image-2.1 scored 60.28 on Qwen-Image-Bench. For hosted deployments on Alibaba Cloud Model Studio, the qwen-image-2.1-turbo API is priced at CNY 0.1 per image with a 120 RPM rate limit, making it 2.5 times cheaper and allowing six times the request throughput of qwen-image-2.1-pro at CNY 0.25 per image and 20 RPM. Compared to alternatives like the 6B Z-Image-Turbo under Apache 2.0 or Black Forest Labs' 9B FLUX.2 klein requiring around 29 GB VRAM, Qwen-Image-2.1-Turbo is distributed under the Qwen Research License, meaning self-hosters must acquire distinct authorization for commercial usage.
This is our own summary of reporting by MarkTechPost



