Qwen-Image-2.1
7B visual generator (32 single-stream DiT layers). Diffusers pipeline QwenImage21Pipeline (install diffusers from git); day-0 ComfyUI, vLLM-Omni, SGLang support. Qwen Research License, check terms before commercial use.
- Input
- text, image
- Output
- image
- License
- qwen-research
- Verified
- 2026-09-30
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | Qwen/Qwen-Image-2.1 | huggingface.co/Qwen/Qwen-Image-2.1 | — |
| ModelScope | — | www.modelscope.cn/models/Qwen/Qwen-Image-2.1 | — |
| GitHub | — | github.com/QwenLM/Qwen-Image-2.1 | — |
| Hugging Face Space (demo) | — | huggingface.co/spaces/Qwen/Qwen-Image-2.1 | — |
Notable capabilities (2)
- Native RGBA generation and editing: Generates transparent images, edits transparent layers and extracts subjects from photos in one model. source
- Multi-reference editing: Up to 10 reference images; local edits via circles, painted annotations or masks, with identity preservation. source
Quick start (from the README):
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16).to("cuda")
image = pipe(prompt="A neon shop sign that reads \"QWEN IMAGE 2.1\"", num_inference_steps=40).images[0]
Timeline entry
- Alibaba releases Qwen-Image-2.1, a 7B open-weights image generation and editing model with native transparency ★★★
On Sept 20, 2026 Alibaba's Qwen team released Qwen-Image-2.1 with open weights: a unified text-to-image and image-editing model whose visual generator has just 7B parameters (32 single-stream DiT layers). It natively generates and edits transparent RGBA layers, takes up to 10 reference images and…
Other Alibaba (Qwen) models
Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-Realtime (Plus) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)