Skip to main content

Qwen-Image Generation & Try-On

Qwen-Image 3.0 is Alibaba's hosted image model on DashScope / Model Studio. OpenTryOn integrates it via QwenImageAdapter for text-to-image, image editing (I2I, 1–3 refs), and virtual try-on (person + garment composition).

This is composition I2I, not Alibaba's dedicated try-on. Dedicated OutfitAnyone-Plus is --model outfitanyone-plus (aitryon-plus, Beijing region). See OutfitAnyone-Plus.

Qwen3.8 (understand) and Qwen-Image (generate / edit / vton) share DASHSCOPE_API_KEY but use different endpoints:

TaskCLIAdapterEndpoint
Understandopentryon understand --model qwen3.8-maxQwenUnderstandAdapterOpenAI-compatible chat (docs)
Generateopentryon generate --model qwen-imageQwenImageAdapter/api/v1/.../multimodal-generation
Editopentryon edit --model qwen-imageQwenImageAdaptersame image API, I2I
VTONopentryon vton --model qwen-imageQwenImageAdapter.generate_virtual_tryonI2I with person + garment
Local generate / edit / VTONopentryon … --model qwen-image-localQwenImageLocalAdapterDiffusers (local docs)

For local Qwen3.8-27B understanding, see Qwen3.8 local. For local Qwen-Image T2I / edit / VTON, see Qwen-Image local (--model qwen-image-local).

Understand → generate / VTON workflow​

Qwen3.8 captions; Qwen-Image paints. Typical fashion loop:

# 1. Describe the garment (Qwen3.8-Max)
opentryon understand --model qwen3.8-max \
--image garment.jpg \
--prompt "Describe this garment: category, color, fabric, cut, and notable details."

# 2a. Text-to-image from that description
opentryon generate --model qwen-image \
--prompt "editorial lookbook photo of a model wearing <paste description>"

# 2b. Or compose the garment onto a person photo
opentryon vton --model qwen-image \
--person-image model.jpg --garment-image garment.jpg \
--garment-description "<paste description>"

VTON here is multi-image composition, not a dedicated garment-fit model. Prefer FLUX VTO or FASHN when drape/fit accuracy matters more than staying on one DashScope key.

Prerequisites​

  1. Alibaba Cloud Model Studio account and API key
  2. Set DASHSCOPE_API_KEY (same key as Wan video and Qwen3.8-Max)
  3. Optional region override via QWEN_IMAGE_BASE_URL
DASHSCOPE_API_KEY=your_dashscope_api_key
# International (default):
# QWEN_IMAGE_BASE_URL=https://dashscope-intl.aliyuncs.com/api/v1
# China:
# QWEN_IMAGE_BASE_URL=https://dashscope.aliyuncs.com/api/v1

Models​

--model-versionRole
qwen-image-3.0-pro (default)Flagship T2I + I2I
qwen-image-3.0Standard 3.0 (quality / speed balance)
qwen-image-2.0-pro / qwen-image-2.0Previous generation

Thinking and prompt rewriting are on by default. --no-thinking and --no-prompt-extend disable them. --prompt-extend-mode agent is T2I-only.

--size is width*height (e.g. 1024*1024). If omitted, 3.0 auto-picks a resolution from the prompt. --n is 1–6.

CLI​

opentryon generate --model qwen-image \
--prompt "editorial lookbook, linen trench on a sunlit terrace" \
--size 1024*1024

opentryon edit --model qwen-image \
--images person.jpg \
--prompt "Change the jacket to black leather, keep the face and pose."

opentryon vton --model qwen-image \
--person-image model.jpg --garment-image garment.jpg \
--garment-description "olive green bomber jacket"

MCP​

Same registry models appear as MCP tools (no extra wiring):

  • generate_qwen_image — T2I (DASHSCOPE_API_KEY)
  • edit_qwen_image — I2I, 1–3 images
  • vton_qwen_image — person + garment composition

Pair with understand_qwen3_8_max for the caption → generate / try-on loop.

Python​

from tryon.api import QwenImageAdapter

adapter = QwenImageAdapter() # qwen-image-3.0-pro by default

images = adapter.generate_text_to_image(
"editorial lookbook, linen trench on a sunlit terrace",
size="1024*1024",
)

edited = adapter.generate_image_edit(
"person.jpg",
prompt="Change the jacket to black leather, keep the face and pose.",
)

tryon = adapter.generate_virtual_tryon(
person="model.jpg",
garment="garment.jpg",
garment_description="olive green bomber jacket",
)
tryon[0].save("qwen_tryon.png")

API Reference​

QwenImageAdapter​

class QwenImageAdapter:
def __init__(
self,
api_key: Optional[str] = None, # DASHSCOPE_API_KEY
model: str = "qwen-image-3.0-pro",
base_url: Optional[str] = None, # QWEN_IMAGE_BASE_URL or intl /api/v1
timeout: float = 300.0,
)

Methods​

  • generate_text_to_image(prompt, size=None, n=1, ...) → List[Image.Image]
  • generate_image_edit(image, prompt, ...) — one image or a list of 1–3
  • generate_multi_image(images, prompt, ...) — I2I composition
  • generate_virtual_tryon(person, garment, prompt=None, garment_description=None, ...)
  • build_tryon_prompt(prompt=None, garment_description=None)

Shared kwargs: negative_prompt, prompt_extend, prompt_extend_mode (direct / agent), enable_thinking, watermark, seed, model.

References​