Skip to main content

Local Models Overview

OpenTryOn provides adapters for local inference models that run directly on your hardware. Unlike the cloud API adapters in tryon.api, these models require local GPU resources but offer several advantages:

  • No API costs: Run unlimited inferences without per-request charges
  • Privacy: Your data never leaves your machine
  • Customization: Fine-tune models for your specific use case
  • Offline capability: Work without internet connectivity

Adding a new local or API model? Start with Model Integration Guidelines.

Available Models​

ModelTypeVRAM RequiredSpeedUse Case
FLUX.2-dev TurboImage Generation12GB+6x fasterFast text-to-image, image-to-image
Kimi-VLImage/Video Understanding24GB+-Open-weight counterpart to the Kimi K2.6/K2.7 Code APIs
Qwen3.8Image/Video Understanding~50GB+ (bf16)-Open Qwen3.8-27B: multimodal understand + thinking; counterpart to DashScope Max
Hy4 previewText LLM (vLLM/SGLang)Datacenter (FP8 ~770GB+, TP=8)-Open tencent/Hy4-preview; OpenTryOn calls localhost OpenAI API — not in-process
Qwen-ImageImage Generation / Edit / VTON~40GB+ (bf16; offload default)-Open Qwen-Image-2512 T2I + Edit-2511 I2I; counterpart to DashScope qwen-image
LTX-2.5Video Generation16GB+ (24GB+ preferred)Distilled few-stepLocal T2V / I2V with synced audio
MiniMax H3Video Generation80GB+ preferred (offload; ~75GB host RAM if int8)HeavyLocal T2V / I2V with stereo audio (768p base)
Wan 2.2Video Generation~12GB+ (TI2V-5B)ModerateLocal T2V / I2V open weights (Wan 3.0 is API-only)
LeffaVirtual Try-On12GB+ recommended~6s on A100Dedicated local VTON (CVPR 2025; MIT code)
CatVTONVirtual Try-On<8GB @ 1024×768ModerateConcatenation VTON (ICLR 2025; CC BY-NC-SA)
Ternary Bonsai 2 27BImage/Text Understanding5.9-8.5GB on disk (ternary-quantized)Fast (CPU/Metal-friendly)OpenAI-compatible client for PrismML's own llama.cpp/MLX server — no torch needed on the OpenTryOn side

Requirements​

Most local models require:

  • CUDA-capable GPU (NVIDIA recommended)
  • PyTorch 2.1+ with CUDA support
  • diffusers >= 0.29.0
  • transformers

Exception: Ternary Bonsai 2 27B and Hy4 preview are OpenAI-compatible HTTP clients for a server you run yourself (locally or on your own cluster) — they need no opentryon[local] extra or GPU on the machine running OpenTryOn itself.

  • accelerate

VRAM Considerations​

Local models are memory-intensive. FLUX.2-dev Turbo supports automatic model selection based on available VRAM:

Available VRAMModel Selection
≥64GBFull precision model
≥48GB8-bit quantized
≥38GB4-bit quantized
<38GB4-bit quantized (with warnings)

Quick Start​

from tryon.models import Flux2TurboAdapter

# Initialize (auto-selects model based on VRAM)
adapter = Flux2TurboAdapter()

# Generate image
images = adapter.generate_text_to_image(
prompt="A fashion model wearing an elegant dress",
width=1024,
height=1024
)
images[0].save("output.png")

Installation​

Install the required dependencies:

pip install diffusers>=0.29.0 transformers accelerate torch

# For quantized models (lower VRAM requirements)
pip install bitsandbytes

Memory Optimization Tips​

  1. Enable CPU Offloading: For GPUs with limited VRAM

    adapter = Flux2TurboAdapter(enable_cpu_offload=True)
  2. Use Attention Slicing: Reduces peak memory at slight speed cost

    adapter = Flux2TurboAdapter(enable_attention_slicing=True)
  3. Lower Resolution: Start with smaller images for testing

    images = adapter.generate_text_to_image(prompt="...", width=512, height=512)
  4. Clear CUDA Cache: Between generations

    import torch
    torch.cuda.empty_cache()