Skip to main content

Qwen3.8-Max Understanding

Qwen3.8-Max is Alibaba's hosted flagship multimodal model on DashScope / Model Studio. OpenTryOn integrates it via QwenUnderstandAdapter for image and video understanding over the OpenAI-compatible Chat Completions API.

The same adapter also serves Qwen3.8-Omni-Flash (--model qwen3.8-omni-flash), Alibaba's native omni-modal model — the only model in this family that accepts audio input (--audio), on top of text/image/video. Same DASHSCOPE_API_KEY, 1M-token context, text-only output.

For local/GPU deployment, see the open-weight Qwen3.8-27B local model.

Capabilities (Qwen3.8 series)​

Qwen3.8 is a native multimodal generation (text + image + video → text), designed for coding, professional work, research, and long-horizon agents — not a separate “VL-only” model ID.

CapabilityHosted Max (qwen3.8-max)Notes
ModalitiesText, image, video inText out
ContextUp to ~1M tokensLarge docs / long multimodal sessions
VideoUp to ~2 hours / 2GB (vendor limits)Many images per request also supported
ThinkingOn by defaultToggle with enable_thinking / --no-thinking
Reasoning depthreasoning_effort: xhigh (default), medium, lowTrades thoroughness vs latency/cost
Coding & agentsStrong long-horizon coding / tool use (vendor)OpenTryOn exposes understand; use chat() or DashScope for full agent loops
Structured output / toolsSupported on DashScopeFunction calling, built-in tools (search / code exec) on the hosted API

OpenTryOn surface today: understand (image and/or video + prompt), plus chat() as an escape hatch for multi-turn / tools. Image generation and virtual try-on use the sibling Qwen-Image adapter (Qwen-Image docs) — same DASHSCOPE_API_KEY.

Family lineup (vendor):

VariantRole
Qwen3.8-MaxHosted MoE flagship (~2.4T total / ~95B active) — CLI qwen3.8-max
Qwen3.8-Omni-FlashHosted native omni-modal (text/image/audio/video in, text out) — CLI qwen3.8-omni-flash
Qwen3.8-27BDense open weights — CLI qwen3.8 (local docs)
Qwen3.8-2.4T-A95BOpen MoE closest to Max — cluster / vLLM–SGLang only

Prerequisites​

  1. Alibaba Cloud Model Studio account and API key
  2. Set DASHSCOPE_API_KEY (same key used for Wan video API)
  3. Optional region override via QWEN_BASE_URL
DASHSCOPE_API_KEY=your_dashscope_api_key
# International (default):
# QWEN_BASE_URL=https://dashscope-intl.aliyuncs.com/compatible-mode/v1
# China:
# QWEN_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1

Installation​

Uses the openai package (already a core opentryon dependency). No extra install step.

Quick Start​

Image Understanding​

from tryon.api import QwenUnderstandAdapter

adapter = QwenUnderstandAdapter() # qwen3.8-max by default

result = adapter.understand_image(
"garment.jpg",
prompt="Describe this outfit: color, pattern, style, fit, and material.",
)
print(result["text"])
print(result["reasoning"]) # thinking content when enabled

Video Understanding​

result = adapter.understand_video(
"runway_clip.mp4",
prompt="Summarize the styling and garments shown in this video.",
)
print(result["text"])

Public https:// media URLs are passed through; local files are inlined as base64 data URIs.

Audio Understanding (Omni-Flash only)​

omni = QwenUnderstandAdapter(model="qwen3.8-omni-flash")

result = omni.understand(
audio="voice_note.wav",
prompt="Transcribe and summarize what is being said.",
)
print(result["text"])

qwen3.8-max raises ValueError if you pass audio — use qwen3.8-omni-flash.

CLI​

opentryon understand --model qwen3.8-max \
--image garment.jpg --prompt "Describe this outfit."

opentryon understand --model qwen3.8-max \
--video runway_clip.mp4 --prompt "Summarize the styling." \
--reasoning-effort medium

opentryon understand --model qwen3.8-max \
--image garment.jpg --no-thinking

opentryon understand --model qwen3.8-omni-flash \
--audio voice_note.wav --prompt "What is being said?"

MCP​

Same registry models appear as MCP tools (no extra wiring):

  • understand_qwen3_8_max — DashScope API (DASHSCOPE_API_KEY)
  • understand_qwen3_8_omni_flash — DashScope API, adds --audio (DASHSCOPE_API_KEY)
  • understand_qwen3_8 — local 27B (local docs)

Caption → generate / try-on (same key): generate_qwen_image, edit_qwen_image, vton_qwen_image — see Qwen-Image.

See MCP Server and mcp-server/README.md.

API Reference​

QwenUnderstandAdapter​

class QwenUnderstandAdapter:
def __init__(
self,
api_key: Optional[str] = None, # DASHSCOPE_API_KEY
model: str = "qwen3.8-max",
base_url: Optional[str] = None, # QWEN_BASE_URL or intl default
)

Methods​

  • understand_image(image, prompt, enable_thinking=None, reasoning_effort=None, ...)
  • understand_video(video, prompt, ...)
  • understand(image=None, video=None, audio=None, prompt=..., ...) — CLI / MCP entry point; audio requires model="qwen3.8-omni-flash"
  • chat(messages, ...) — multi-turn / tools escape hatch (raw ChatCompletion)

Return dict keys: text, reasoning, model, usage.

References​