Skip to main content

CatVTON (open-weight VTON)

CatVTON (ICLR 2025) concatenates the person and garment in latent space on SD 1.5 inpainting, skipping text cross-attention. The trainable attention adapters are small (~50M); full inference is about 899M params and typically under 8GB VRAM at 1024×768.

Registry idcatvton
Weightszhengchong/CatVTON
Codegithub.com/Zheng-Chong/CatVTON
PaperarXiv:2407.15886
VRAM<8GB @ 1024×768 (fp16/bf16)
LicenseCC BY-NC-SA 4.0 (code + checkpoints) — not for commercial D2C

The same HF repo also contains a FLUX.1-Fill LoRA (flux-lora/, 37.4M). Official FLUX inference code was not released with the LoRA; OpenTryOn implements the documented SD 1.5 pipeline (mix / vitonhd / dresscode attention folders).

Install​

pip install opentryon[local]

The SD 1.5 inpainting base (runwayml/stable-diffusion-inpainting) may be gated. Override with a community mirror:

export CATVTON_BASE_MODEL=botp/stable-diffusion-v1-5-inpainting

Environment​

# CATVTON_ATTN_CKPT=zhengchong/CatVTON
# CATVTON_BASE_MODEL=runwayml/stable-diffusion-inpainting
# HF_TOKEN=hf_...

CLI​

opentryon vton --model catvton \
--person-image person.jpg --garment-image dress.jpg \
--garment-type dresses --attn-version mix --dry-run

opentryon vton --model catvton \
--person-image person.jpg --garment-image top.jpeg \
--mask-image agnostic.png --steps 50 --width 768 --height 1024

--attn-version mix is the 1024 mix checkpoint (default). vitonhd and dresscode are the 512 dataset-specific adapters.

Pass --mask-image (white = replace) when you have an agnostic mask. Without it, OpenTryOn uses a geometric upper/lower/dress rectangle.

Python​

from tryon.models import CatVTONAdapter

adapter = CatVTONAdapter(attn_version="mix")
images = adapter.generate_and_decode("person.jpg", "garment.jpg")
images[0].save("tryon.png")

MCP​

Tool: vton_catvton. Needs pip install opentryon[local] + GPU. Restart MCP so TryOn Studio lists it.

See also​

  • Leffa — MIT-licensed code, stronger detail/logo story, heavier stack
  • Hosted dedicated VTON if you cannot use CC BY-NC-SA weights in production