Configuration
Learn how to configure OpenTryOn for your specific needs. The same opentryon/.env file is read by the CLI, the MCP server, and TryOn Studio Connect (Studio never stores keys itself).
Environment Variables
OpenTryOn uses environment variables for configuration. Create a .env file in your project root:
Preprocessing (Required for Local Preprocessing)
# U2Net Model Checkpoints (Required for garment/human segmentation)
U2NET_CLOTH_SEG_CHECKPOINT_PATH=path/to/cloth_segm.pth
U2NET_SEGM_CHECKPOINT_PATH=path/to/u2net.pth
# Optional: GPU Configuration
CUDA_VISIBLE_DEVICES=0
# Optional: Logging
LOG_LEVEL=INFO
API Integrations (Optional - Only configure APIs you plan to use)
# Segmind Try-On Diffusion API
SEGMIND_API_KEY=your_segmind_api_key
# Kling AI Virtual Try-On API
KLING_AI_API_KEY=your_kling_api_key
KLING_AI_SECRET_KEY=your_kling_secret_key
KLING_AI_BASE_URL=https://api-singapore.klingai.com # Optional
# Amazon Nova Canvas (AWS Bedrock)
AWS_ACCESS_KEY_ID=your_aws_access_key
AWS_SECRET_ACCESS_KEY=your_aws_secret_key
AMAZON_NOVA_REGION=us-east-1 # Options: us-east-1, ap-northeast-1, eu-west-1
AMAZON_NOVA_MODEL_ID=amazon.nova-canvas-v1:0 # Optional
# Google Gemini (Nano Banana Image Generation)
GEMINI_API_KEY=your_gemini_api_key
# Google Vertex Virtual Try-On (virtual-try-on-001) — not GEMINI_API_KEY
GOOGLE_CLOUD_PROJECT=your_gcp_project_id
# GOOGLE_CLOUD_LOCATION=global
# BFL AI (FLUX.2 Image Generation)
BFL_API_KEY=your_bfl_api_key
# Moonshot AI (Kimi K2.6 / K2.7 Code multimodal understanding)
MOONSHOT_API_KEY=your_moonshot_api_key
# Tencent TokenHub (Hy4 preview LLM) — not Moonshot / DashScope
TOKENHUB_API_KEY=your_tokenhub_api_key
# TOKENHUB_BASE_URL=https://tokenhub-intl.tencentcloudmaas.com/v1
# Alibaba DashScope (Wan, Qwen3.8-Max, Qwen3.8-Omni-Flash, Qwen-Image, OutfitAnyone-Plus)
DASHSCOPE_API_KEY=your_dashscope_api_key
# Z.ai / Zhipu (GLM-5.3-FlashX multimodal understanding)
ZAI_API_KEY=your_zai_api_key
# ZAI_BASE_URL=https://api.z.ai/api/paas/v4
# Photoroom Virtual Try-On + Virtual Model
PHOTOROOM_API_KEY=your_photoroom_api_key
# MiniMax Hailuo 2.3 + MiniMax H3 / H3 Max video (same key; H3 uses V2)
MINIMAX_API_KEY=your_minimax_api_key
# Fal (third-party MiniMax H3 Max — T2V / I2V / R2V)
FAL_KEY=your_fal_key
# NVIDIA NIM (Nemotron Omni understand, Cosmos 3 Reasoner, Cosmos 3 Generator)
NVIDIA_API_KEY=your_nvidia_api_key
# Meta Model API (Muse Image generate/edit/vton)
MODEL_API_KEY=your_meta_model_api_key
# Pruna AI (P-Image, P-Image-Ideogram, P-Image-Edit, try-on, P-Video family incl. P-Video-2-Pro)
PRUNA_API_KEY=your_pruna_api_key
Local server models (Optional — no cloud key)
# Ternary Bonsai 2 27B — you run its llama.cpp/MLX server yourself;
# OpenTryOn is just an OpenAI-compatible client against it.
# BONSAI_BASE_URL=http://127.0.0.1:8080/v1
# BONSAI_API_KEY=... # only if your server enforces auth
Datasets (Optional - Only if using HuggingFace datasets)
# HuggingFace datasets cache (for Subjects200K)
# Defaults to ~/.cache/huggingface/datasets if not set
HF_DATASETS_CACHE=path/to/cache
Planner / Studio chat (optional — only if you use planner_agent)
The cheap intent model is separate from image/VTON/video keys. Studio Agent chat calls MCP planner_agent; set these on the MCP host and restart the server. See Planner Agent and TryOn Studio.
OPENTRYON_AGENT_LLM_PROVIDER=openai
OPENTRYON_PLANNER_LLM_MODEL=gpt-4o-mini
# OPENAI_API_KEY=... # or ANTHROPIC_API_KEY / GEMINI_API_KEY
Note: You only need to configure the APIs and features you plan to use. For example:
- Preprocessing only: Only U2Net checkpoints required
- API integrations only: Only API keys required (no local models needed)
- Datasets only: No configuration needed (automatic download/caching)
Loading Environment Variables
Always load environment variables before using OpenTryOn:
from dotenv import load_dotenv
load_dotenv()
# Now import and use OpenTryOn modules
from tryon.preprocessing import segment_garment
from tryon.api import SegmindVTONAdapter
from tryon.datasets import FashionMNIST
Getting API Keys
Segmind Try-On Diffusion
- Sign up at Segmind API Portal
- Obtain your API key from the dashboard
- Add to
.env:SEGMIND_API_KEY=your_key
Kling AI Virtual Try-On
- Sign up at Kling AI Developer Portal
- Obtain API key (access key) and secret key
- Add to
.env:KLING_AI_API_KEY=your_api_key
KLING_AI_SECRET_KEY=your_secret_key
Amazon Nova Canvas
- Set up AWS account with Bedrock access
- Enable Nova Canvas in AWS Bedrock console (Model access section)
- Configure AWS credentials (via
.envor AWS CLI):AWS_ACCESS_KEY_ID=your_access_key
AWS_SECRET_ACCESS_KEY=your_secret_key
AMAZON_NOVA_REGION=us-east-1
Google Gemini (Nano Banana)
- Sign up at Google AI Studio
- Obtain API key from API Keys page
- Add to
.env:GEMINI_API_KEY=your_key
BFL AI (FLUX.2)
- Sign up at BFL AI
- Obtain your API key from the BFL AI dashboard
- Add to
.env:BFL_API_KEY=your_key
Moonshot AI (Kimi K2.6 / K2.7 Code)
- Sign up at platform.kimi.ai
- Obtain your API key from the API Keys console
- Add to
.env:MOONSHOT_API_KEY=your_key
Tencent TokenHub (Hy4 preview)
- Follow TokenHub Chat Completions and create an API key
- Add to
.env:TOKENHUB_API_KEY=your_key - Optional:
TENCENT_TOKENHUB_API_KEY(alias) orTOKENHUB_BASE_URL(default international endpoint) - Local weights twin (
hy4-preview-local) does not use this key — serve vLLM/SGLang and setHY4_BASE_URL
See Hy4 TokenHub and Hy4 local.
Alibaba DashScope (Qwen3.8, Qwen-Image, Wan)
-
Sign up at Alibaba Cloud Model Studio
-
Create an API key for your region
-
Add to
.env:DASHSCOPE_API_KEY=your_key
# Optional: QWEN_BASE_URL for Qwen3.8-Max chat (OpenAI-compatible)
# Optional: QWEN_IMAGE_BASE_URL for Qwen-Image T2I / I2I / VTONSame key covers
understand --model qwen3.8-max,understand --model qwen3.8-omni-flash(adds--audio),generate|edit|vton --model qwen-image,video-generate --model wan-api/wan-3.0, and Beijing-regionvton --model outfitanyone-plus(aitryon-plus). International keys used for Qwen/Wan do not unlock OutfitAnyone-Plus.Local open-weight twin (
pip install opentryon[local], CUDA, recent Diffusers):# Optional overrides; defaults are the official HF snapshots
# QWEN_IMAGE_LOCAL_MODEL_ID=Qwen/Qwen-Image-2512
# QWEN_IMAGE_EDIT_MODEL_ID=Qwen/Qwen-Image-Edit-2511
# QWEN_IMAGE_LOCAL_PATH=/path/to/local/t2i-snapshot
# QWEN_IMAGE_EDIT_PATH=/path/to/local/edit-snapshotCLI:
opentryon generate|edit|vton --model qwen-image-local. See Qwen-Image local.
Local dedicated VTON (Leffa / CatVTON)
No API key. Needs pip install opentryon[local] and a CUDA GPU.
vton --model leffa— Leffa. OptionalLEFFA_HOME/LEFFA_CKPT.vton --model catvton— CatVTON (CC BY-NC-SA 4.0). OptionalCATVTON_BASE_MODELif the SD 1.5 inpainting repo is gated.
Photoroom (Virtual Try-On / Virtual Model)
-
Activate the API at app.photoroom.com/api
-
Add to
.env:PHOTOROOM_API_KEY=your_key -
Optional watermarked tests: prefix the key with
sandbox_or setPHOTOROOM_SANDBOX=1Covers
vton --model photoroom-vton(shopper photo + product) andvton --model photoroom-virtual-model(flat-lay → on-model). Plus / Enterprise Image Editing API. See Photoroom.
MiniMax (Hailuo 2.3 + H3 + H3 Max)
-
Sign up at MiniMax Open Platform
-
Create an interface key from API keys
-
Add to
.env:MINIMAX_API_KEY=your_keySame key covers
video-generate --model hailuo-2.3(V1),video-generate --model minimax-h3(V2 H3), andvideo-generate --model minimax-h3-max(V2 H3 Max, fast). H3 on the API is billed as pay-as-you-go video.Local open-weight twin (
pip install opentryon[local], CUDA, Diffusers from main):--model minimax-h3-local. The Community License for those weights excludes US/EU/UK/South Korea unless separately authorized. See MiniMax H3 local.
Fal (MiniMax H3 Max)
-
Create a key at Fal API keys
-
Add to
.env:FAL_KEY=your_key(FAL_API_KEYis an alias)Covers
video-generate --model fal-h3-max(T2V / I2V / R2V) andvideo-generate --model fal-h3-max-lipsync(portrait + audio → lip-synced video). This is a third-party hoster, not MiniMax’s V2 API. First-party Max remains--model minimax-h3-max. See MiniMax H3 Max (Fal).
Z.ai / Zhipu (GLM-5.3-FlashX)
-
Create a key at Z.ai / the Z.ai API docs
-
Add to
.env:ZAI_API_KEY=your_keyCovers
understand --model glm-5.3-flashx(text/image/video understanding, 200 tok/s serving tier of GLM-5.3-Flash). See GLM-5.3-FlashX.
Ternary Bonsai 2 27B (local server, no cloud key)
No API key. Start the model's own llama.cpp (PrismML fork, prism-b10658+) or MLX server yourself, then point OpenTryOn at it:
# BONSAI_BASE_URL=http://127.0.0.1:8080/v1 # default
# BONSAI_API_KEY=... # only if your server enforces auth
CLI: opentryon understand --model ternary-bonsai-2-27b. See Ternary Bonsai 2 27B.
NVIDIA NIM (Nemotron / Cosmos)
-
Create a key at build.nvidia.com
-
Add to
.env:NVIDIA_API_KEY=your_keySame key covers
understand --model nemotron-omni,understand --model cosmos3-reasoner, andvideo-generate --model cosmos3. OptionalCOSMOS3_INFER_URLpoints a self-hosted Generator NIM athttp://127.0.0.1:8000/v1/infer. See NVIDIA NIM.
Meta Model API (Muse Image)
-
Create a key in the Model API dashboard
-
Add to
.env:MODEL_API_KEY=your_key(aliases:META_MODEL_API_KEY,MUSE_API_KEY)Covers
generate|edit|vton --model muse-image. Muse Video has no developer API or open weights yet.
Configuration Options
GPU Configuration
Specify which GPU to use:
import os
os.environ["CUDA_VISIBLE_DEVICES"] = "0" # Use first GPU
Model Checkpoint Paths
Set custom checkpoint paths:
import os
os.environ["U2NET_CLOTH_SEG_CHECKPOINT_PATH"] = "/custom/path/cloth_segm.pth"
os.environ["U2NET_SEGM_CHECKPOINT_PATH"] = "/custom/path/u2net.pth"
Logging Configuration
Configure logging level:
import logging
logging.basicConfig(level=logging.INFO)
Default Settings
OpenTryOn uses sensible defaults:
- Image Size: Automatically resized based on model requirements
- Batch Size: 1 (can be adjusted for batch processing)
- Device: Auto-detects CUDA if available, falls back to CPU
- Normalization: Images normalized to [-1, 1] range
Custom Configuration
You can override defaults when calling functions:
from tryon.preprocessing.extract_garment_new import extract_garment
from PIL import Image
import torch
# Use specific device
device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu")
# Load model once for efficiency
net = load_cloth_segm_model(device, os.environ.get("U2NET_CLOTH_SEGM_CHECKPOINT_PATH"))
# Use pre-loaded model
image = Image.open("garment.jpg")
garments = extract_garment(
image=image,
cls="upper",
resize_to_width=400,
net=net, # Reuse model
device=device
)
Quick Configuration Examples
Preprocessing Only
U2NET_CLOTH_SEG_CHECKPOINT_PATH=./models/cloth_segm.pth
U2NET_SEGM_CHECKPOINT_PATH=./models/u2net.pth
API Integrations Only (No Local Models)
SEGMIND_API_KEY=your_segmind_key
GEMINI_API_KEY=your_gemini_key
BFL_API_KEY=your_bfl_key
Full Setup (Preprocessing + APIs + Datasets)
# Preprocessing
U2NET_CLOTH_SEG_CHECKPOINT_PATH=./models/cloth_segm.pth
U2NET_SEGM_CHECKPOINT_PATH=./models/u2net.pth
# APIs
SEGMIND_API_KEY=your_segmind_key
KLING_AI_API_KEY=your_kling_key
KLING_AI_SECRET_KEY=your_kling_secret
GEMINI_API_KEY=your_gemini_key
BFL_API_KEY=your_bfl_key
AWS_ACCESS_KEY_ID=your_aws_key
AWS_SECRET_ACCESS_KEY=your_aws_secret
AMAZON_NOVA_REGION=us-east-1
Best Practices
- Always use
.envfile: Never commit API keys or paths to version control - Load environment variables first: Before importing any OpenTryOn modules
- Use absolute paths: For checkpoint paths to avoid issues
- Check GPU availability: Verify CUDA before running intensive operations
- Only configure what you need: Don't add API keys for services you won't use
- Keep
.envin.gitignore: Protect your credentials
Next Steps
- Quick Start Guide: See examples of using APIs, datasets, and preprocessing
- API Reference: Complete API documentation
- Datasets Module: Learn about available datasets
- Preprocessing: Preprocessing documentation