Local Models Overview
OpenTryOn provides adapters for local inference models that run directly on your hardware. Unlike the cloud API adapters in tryon.api, these models require local GPU resources but offer several advantages:
- No API costs: Run unlimited inferences without per-request charges
- Privacy: Your data never leaves your machine
- Customization: Fine-tune models for your specific use case
- Offline capability: Work without internet connectivity
Adding a new local or API model? Start with Model Integration Guidelines.
Available Models
| Model | Type | VRAM Required | Speed | Use Case |
|---|---|---|---|---|
| FLUX.2-dev Turbo | Image Generation | 12GB+ | 6x faster | Fast text-to-image, image-to-image |
| Kimi-VL | Image/Video Understanding | 24GB+ | - | Open-weight counterpart to the Kimi K2.6/K2.7 Code APIs |
| Qwen3.8 | Image/Video Understanding | ~50GB+ (bf16) | - | Open Qwen3.8-27B: multimodal understand + thinking; counterpart to DashScope Max |
| Hy4 preview | Text LLM (vLLM/SGLang) | Datacenter (FP8 ~770GB+, TP=8) | - | Open tencent/Hy4-preview; OpenTryOn calls localhost OpenAI API — not in-process |
| Qwen-Image | Image Generation / Edit / VTON | ~40GB+ (bf16; offload default) | - | Open Qwen-Image-2512 T2I + Edit-2511 I2I; counterpart to DashScope qwen-image |
| LTX-2.5 | Video Generation | 16GB+ (24GB+ preferred) | Distilled few-step | Local T2V / I2V with synced audio |
| MiniMax H3 | Video Generation | 80GB+ preferred (offload; ~75GB host RAM if int8) | Heavy | Local T2V / I2V with stereo audio (768p base) |
| Wan 2.2 | Video Generation | ~12GB+ (TI2V-5B) | Moderate | Local T2V / I2V open weights (Wan 3.0 is API-only) |
| Leffa | Virtual Try-On | 12GB+ recommended | ~6s on A100 | Dedicated local VTON (CVPR 2025; MIT code) |
| CatVTON | Virtual Try-On | <8GB @ 1024×768 | Moderate | Concatenation VTON (ICLR 2025; CC BY-NC-SA) |
| Ternary Bonsai 2 27B | Image/Text Understanding | 5.9-8.5GB on disk (ternary-quantized) | Fast (CPU/Metal-friendly) | OpenAI-compatible client for PrismML's own llama.cpp/MLX server — no torch needed on the OpenTryOn side |
Requirements
Most local models require:
- CUDA-capable GPU (NVIDIA recommended)
- PyTorch 2.1+ with CUDA support
- diffusers >= 0.29.0
- transformers
Exception: Ternary Bonsai 2 27B and Hy4 preview
are OpenAI-compatible HTTP clients for a server you run yourself (locally or
on your own cluster) — they need no opentryon[local] extra or GPU on the
machine running OpenTryOn itself.
- accelerate
VRAM Considerations
Local models are memory-intensive. FLUX.2-dev Turbo supports automatic model selection based on available VRAM:
| Available VRAM | Model Selection |
|---|---|
| ≥64GB | Full precision model |
| ≥48GB | 8-bit quantized |
| ≥38GB | 4-bit quantized |
| <38GB | 4-bit quantized (with warnings) |
Quick Start
from tryon.models import Flux2TurboAdapter
# Initialize (auto-selects model based on VRAM)
adapter = Flux2TurboAdapter()
# Generate image
images = adapter.generate_text_to_image(
prompt="A fashion model wearing an elegant dress",
width=1024,
height=1024
)
images[0].save("output.png")
Installation
Install the required dependencies:
pip install diffusers>=0.29.0 transformers accelerate torch
# For quantized models (lower VRAM requirements)
pip install bitsandbytes
Memory Optimization Tips
-
Enable CPU Offloading: For GPUs with limited VRAM
adapter = Flux2TurboAdapter(enable_cpu_offload=True) -
Use Attention Slicing: Reduces peak memory at slight speed cost
adapter = Flux2TurboAdapter(enable_attention_slicing=True) -
Lower Resolution: Start with smaller images for testing
images = adapter.generate_text_to_image(prompt="...", width=512, height=512) -
Clear CUDA Cache: Between generations
import torch
torch.cuda.empty_cache()