Skip to main content

Integrate next

Living candidate queue for new adapters. This is not a commitment and it is not the v0.1.0 product roadmap.

FileRole
This pageCanonical list (git + docs site)
ROADMAP.mdProduct slices (train / eval / one local VTON / agents)
.cursor/skills/integrate-model/How to integrate once you pick a row

Surveyed: 29 August 2026 · Sources: build.nvidia.com/models, Google virtual-try-on-001, Alibaba OutfitAnyone-Plus, Photoroom Virtual Try-On, CatVTON, Leffa, Nemotron, Cosmos 3.

How to use: pick a next row → follow the integrate-model skill (Path A first-party API, Path B local). After ship, move the row to Shipped and bump the date.

StatusMeaning
nextWorth integrating when someone is ready
watchInteresting; wait for a clearer API or fashion fit
blockedNo first-party API / weights / license yet
skipOut of scope or we already cover it another way

NVIDIA access is usually NVIDIA_API_KEY on build.nvidia.com (hosted NIM) and/or a self-hosted NIM container. Prefer the hosted NIM for Path A; HF weights + Diffusers/vLLM for Path B. Do not wrap FLUX / Qwen-Image / Ideogram a second time through NIM — we already have first-party adapters.

OpenTryOn services today: vton · generate · edit · understand · video-generate · bg-remove. Rows marked new service need a new CLI/MCP category before an adapter.


Wave 1 — NVIDIA (fits existing services)​

Highest leverage: one NIM provider key unlocks understand + video. Nemotron is not a T2I/VTON family — it is open multimodal understanding / agents. Cosmos is NVIDIA’s generation stack.

CandidateCategoryPathSuggested idWhyStatus
Nemotron 3 Nano Omni 30B-A3B ReasoningUnderstand (image + video + audio + text)A NIMnemotron-omniHosted NIM chat. Path B local weights later.shipped
Cosmos 3 Generator (nano, 8B)Text-to-video, image-to-videoA NIM (POST infer)cosmos3T2V if prompt only; I2V if image set. Optional COSMOS3_INFER_URL.shipped
Cosmos 3 Reasoner (nano)Understand (physical / video)A NIMcosmos3-reasonerWorld-model VLM.shipped
Nemotron Nano 12B V2 VLUnderstand (image + video)A NIM—Predecessor to Omni. Only if Omni is too large.watch
Cosmos 3 Generator super (32B)T2V / I2VA NIM (NIM_MODEL_SIZE=super)cosmos3-superSame API as nano; heavier GPU.watch
Cosmos Predict 2.5 2BT2V / I2V / V2VA NIMcosmos-predict-2.5Older WFM; prefer Cosmos 3 unless Predict-only features are needed.watch
Nemotron 3 Nano / Super / Ultra (text)Text-only agentsA NIM—Coding/planning MoE. No image/video I/O. Planner already uses other LLMs.skip
Nemotron OCR / Parse / page-tableDocument OCRA NIM—Useful later for PDP/catalog agents, not invoke-layer media.watch
Nemotron 3.5 Content SafetySafety classifierA NIM—Guardrail, not a generation/understand tool.watch

Nemotron 3 family (context, Aug 2026)

ModelRoleOpenTryOn fit
Nemotron 3 Nano 30B-A3BText agents, 1M contextLow (no media)
Nemotron 3 Super 120B-A12BLarger text agentsLow
Nemotron 3 Ultra 550B-A55BFlagship textLow
Nemotron 3.5 Lightning 30B-A3BFast text agentsLow
Nemotron 3 Nano Omni 30B-A3BImage / video / speech / textHigh — Wave 1
Nemotron Voicechat / ASR / Magpie TTSSpeechLater (no speech service yet)
Llama-Nemotron embed/rerank VLRAG embeddingsOut of scope

Wave 2 — Virtual try-on (fashion / D2C / marketplace)​

Compile-only as of 29 Aug 2026 — do not integrate until asked. Developers building fitting rooms, PDP/catalog on-model shots, and marketplace listing tools need dedicated person+garment try-on, not another general I2I compose.

Trust order for Path A: hyperscalers and durable public platforms first (Google, Amazon, Alibaba, BFL, Kuaishou, Meta). Fashion specialists we already ship (FASHN, Pruna) stay. Smaller photo-API vendors are watch unless a customer names them. Do not add Fal / Replicate / PiAPI wrappers of models we already call first-party.

NVIDIA has no dedicated VTON NIM. Product roadmap Slice D is still: pick one local OSS path for v0.1.0.

Already in OpenTryOn (do not re-add)​

Registry idVendorKindNotes
flux-vtoBlack Forest LabsDedicated VTON APIFirst-party FLUX VTO
google-vtonGoogle Cloud VertexDedicated VTON APIvirtual-try-on-001; ADC + GOOGLE_CLOUD_PROJECT, not GEMINI_API_KEY
outfitanyone-plusAlibaba Cloud Model StudioDedicated VTON APIaitryon-plus; Beijing-region DASHSCOPE_API_KEY, not Qwen-Image compose
photoroom-vton / photoroom-virtual-modelPhotoroomDedicated VTON + catalog APIImage Editing /v2/edit; shopper try-on or flat-lay → on-model
nova-canvasAmazon BedrockDedicated VTON APIGarment classes incl. footwear
kling-aiKuaishou (Kling / Kolors)Dedicated VTON APIFirst-party Kolors v1 / v1.5
fashn-tryon-max / fashn-tryon-v1.6FASHNDedicated VTON APIFashion suite; v1.6 is the fast e-comm path
p-image-tryonPrunaDedicated VTON APIMulti-garment (up to 11 refs)
segmindSegmindHosted try-on diffusionThird-party hoster; keep, do not add more hosters
nano-banana-2-liteGoogle GeminiComposition I2INot Vertex virtual-try-on-001
qwen-image / qwen-image-localAlibaba QwenComposition I2INot OutfitAnyone aitryon-plus
muse-imageMetaComposition I2IMulti-ref edit, not a garment-fit model

Path A — dedicated VTON APIs to add​

Prefer first-party APIs. Related catalog jobs (product→model, model-swap, parsing) are listed only when they are that vendor’s try-on product.

CandidateVendor durabilityTaskSuggested idWhyStatus
Google virtual-try-on-001Google Cloud (GA 20 Jan 2026; listed discontinue 20 Jan 2027)Shopper / catalog image VTONgoogle-vtonDedicated Vertex / Gemini Enterprise predict API. Person + product image, 1–4 samples, C2PA watermark. Auth is ADC / GCP project, not GEMINI_API_KEY. Distinct from Nano Banana compose. Docsshipped
Alibaba OutfitAnyone-Plus (aitryon-plus)Alibaba Cloud Model StudioImage VTON + combo top/bottomoutfitanyone-plusDedicated DashScope try-on (async). Top, bottoms, dress, face restore, parsing companion aitryon-parsing-v1. Same company as Qwen; Beijing-region key, not the Qwen-Image compose path. Docsshipped
Photoroom Virtual Try-On / Virtual ModelPhotoroom (widely used e-comm photo API)Fitting room or garment→lifestyle modelphotoroom-vton / photoroom-virtual-modelAPI-first catalog/shopper flows. Virtual Model is “flat-lay in, on-model out” (no person photo). Complements dedicated person+SKU VTON. API · Productshipped
Pixelcut Try-OnPixelcut (e-comm photo; Shopify-heavy)Image VTON + garment transferpixelcut-vtonREST /v1/try-on; upper/lower/full. Smaller than Google/Alibaba. APIwatch
FitroomSpecialist startupCombo top+bottom in one requestfitroomStrong e-comm DX; weaker long-term vendor signal.watch
ClaidSpecialistCatalog try-on—Photo-API vendor; overlap with Photoroom.watch
BytePlus Effects / live AR try-onByteDanceReal-time AR, not still VTON—SDK/effects stack. Seedream image/edit is already in the registry (seedream).watch
Tencent Cloud FitDiT (hosted)TencentCommercial FitDiT—Open weights are NC; Tencent Cloud is the commercial door. Confirm a public REST API before Path A.watch
Adobe Firefly ServicesAdobeGeneral gen/edit—Commercially durable Creative Cloud; no dedicated person+garment VTON API found.skip
Shopify / Google Shopping / Walmart ZeekitPlatform lock-inIn-app try-on—Not a developer API we can register.skip
Kling via Fal / PiAPI / ReplicateAggregatorsSame Kolors VTON—We already have first-party kling-ai.skip
Snap Camera Kit / glasses ARSnapAccessory AR—Different modality (mesh/AR), not image VTON.watch

Related try-on jobs (same developers, not a second vton clone):

JobWhat they callPrefer
Shopper fitting roomPerson selfie + SKU photoGoogle VTO, FASHN, Kling, BFL, Amazon, OutfitAnyone
Catalog on-modelFlat-lay → generated modelPhotoroom Virtual Model; FASHN product-to-model (vendor suite — do not invent a parallel adapter until asked)
Multi-SKU outfitTop + bottoms one callOutfitAnyone combo, Fitroom combo, Pruna multi-ref (shipped)
Model swap / consistent modelFace/body swap, keep garmentFASHN model-swap (vendor suite)
Video try-onTemporal garment on a clipCatV2TON local; no durable first-party video-VTON API picked yet
Parsing / hotspotsGarment masks, bboxesAlibaba aitryon-parsing-v1; OpenTryOn already has preprocess helpers

Path B — local / open-weight VTON​

Pick one for v0.1.0 Slice D (tryon.models + opentryon[local]). Many research checkpoints are CC BY-NC-SA — fine for OSS demos, a problem for D2C/marketplace production. Confirm license before making one the default.

CandidateOriginSuggested idVRAM / notesLicense (typical)Status
LeffaCVPR 2025; HF franciszzj/LeffaleffaDiffusers; VITON-HD + DressCode try-on + pose transfer; strong detail/logo storyCode MIT; confirm weight card for commercial D2Cshipped
CatVTON + CatVTON-FLUX LoRAICLR 2025; FLUX.1-Fill LoRA ~37Mcatvton<8GB @ 1024×768; SD 1.5 concatenation pipeline (FLUX LoRA weights exist; official FLUX infer code not released)CC BY-NC-SA 4.0 (code + checkpoints); FLUX-Fill base has its own termsshipped
IDM-VTONECCV 2024; yisol/IDM-VTONidm-vtonHigher fidelity; ~18–24GB typicalCC BY-NC-SAwatch
OOTDiffusionlevihsu/OOTDiffusionootdiffusionCommunity baseline; setup scripts already under tryon/Check repowatch
FitDiTTencent-affiliated DiT; BoyuanJiang/FitDiTfitditHigh garment-detail DiT; ComfyUI existsCC BY-NC-SA; commercial via Tencent Cloudwatch
CatV2TONSame lab as CatVTONcatv2tonVideo try-on; needs a video-VTON design passCheck repo (likely NC like CatVTON)watch
FLUX-fill LoRA (train slice)BFL Fill + brand LoRA—Not a third cloud VTON; opentryon train pathFLUX termswatch
OutfitAnyone weightsHumanAIGC / Alibaba paper—Demos lock person upload; use aitryon-plus API insteadRestricted demosskip
VITON-HD / StableVITON2022–2023 warping/diffusion—Superseded for new workMixedskip
Qwen-Image-Edit-2511 localAlready qwen-image-local—Composition I2I, not a VTON specialist—shipped

By capability (NVIDIA + fashion-relevant others)​

Text-to-image​

CandidatePathNotesStatus
SANA-Sprint (NVLabs, HF Diffusers)BFast local T2I (1–4 steps). Not a NIM.next
Cosmos 3 T2IB / vLLM-OmniNIM Generator does not expose T2I (one-frame video only in TRT-LLM).watch
FLUX.1-dev / schnell / Kontext via NIMAWe already have BFL FLUX.2. Do not add a second FLUX stack.skip
FLUX.2 Klein 4B via NIMADistilled FLUX.2; only if BFL does not offer Klein.watch
SD 3.5 Large via NIMACrowded T2I table; low fashion differentiation.skip
Qwen-Image via NIMAAlready qwen-image (DashScope) + local.skip
NVIDIA Edify (Getty/Shutterstock)—NIM preview retired 6 June 2025.skip

Image-to-image / edit​

CandidatePathNotesStatus
FLUX.1 Kontext via NIMAIn-context edit. Skip unless we want NIM as a fallback host.skip
Cosmos Transfer 2.5 2BA NIMVideo control transfer (edge/depth/seg/vis), not still I2I. See video.—
SANA-Sprint ControlNetBLocal realtime I2I if we take SANA.watch

Virtual try-on​

NVIDIA / Nemotron has no VTON NIM. Cloud dedicated VTON is already broad (FLUX VTO, Google Vertex, OutfitAnyone-Plus, Photoroom, Amazon Nova, Kling, FASHN, Pruna). Local weights shipped: Leffa (leffa, MIT code) and CatVTON (catvton, CC BY-NC-SA). Full tables: Wave 2 above.

Text-to-video​

CandidatePathNotesStatus
Cosmos 3 Generator nanoAWave 1 shipped (cosmos3). Physics-aware; fashion lookbooks are a stretch but the API is clean.shipped
Cosmos Predict 2.5ALegacy WFM.watch
Muse Video—Consumer preview only; no Meta Model API / weights.blocked

Audio-to-video / talking head​

New CLI service (e.g. audio-to-video) — do not stuff this into video-generate without a design pass.

CandidatePathNotesStatus
NVIDIA LipSync NIMAAudio → lip-dubbed video. Closest official A2V on the NIM catalog.next
Cosmos 3 + soundB / vLLM-OmniT2V/I2V with synchronized audio. Not on Generator NIM.watch
Pruna P-Video-Avatar—Already shipped (p-video-avatar).shipped

Image & video understanding / multimodal​

CandidatePathNotesStatus
Nemotron 3 Nano OmniAWave 1 shipped. Native audio (unlike current Kimi/Qwen understand tools).shipped
Cosmos 3 Reasoner / Reason2 8BAReasoner shipped (cosmos3-reasoner). Reason2 remains watch.shipped / watch
Muse Glimmer 30B (on NIM)AMeta multimodal on NVIDIA’s catalog; we already have first-party Muse Image. Glimmer is understand, not gen.watch
Llama 3.2 11B/90B Vision on NIMAOlder VLMs; Omni supersedes for new work.skip

3D model generation​

New CLI service (e.g. generate-3d). Roadmap lists 3D VTON under Later.

CandidatePathNotesStatus
Microsoft TRELLIS NIMAText-to-3D and image-to-3D meshes; active NVIDIA 3D NIM (Edify 3D is gone).next
TRELLIS.2 4BBLocal; Windows-skewed wheels as of 2026 — verify Linux before Path B.watch
Video VTON / 3D VTON—Research; no product API picked.watch

Out of scope for this list​

Biology (AlphaFold, Evo2), CFD, weather, routing, chip sim, protein design — on the NIM catalog, not fashion media.


Suggested integration order​

  1. nemotron-omni / cosmos3 / cosmos3-reasoner shipped (Path A, 29 Aug 2026).
  2. google-vton shipped (Path A Vertex virtual-try-on-001, 29 Aug 2026).
  3. outfitanyone-plus / photoroom-vton / photoroom-virtual-model shipped (Path A, 29 Aug 2026).
  4. leffa / catvton shipped (Path B local VTON, 1 Sep 2026).
  5. New services only after one local VTON: LipSync (A2V), TRELLIS (3D).
  6. SANA-Sprint if we want a fast local T2I that is not another FLUX/Qwen clone.
  7. Optional: nemotron-omni-local if someone will run 30B-A3B.

Shipped (do not re-add)​

Invoke-layer highlights already in the registry: FLUX.2 (+ Turbo local), Nano Banana family, GPT Image (1.5 + ChatGPT Images 2.5 Flare/Sunburst), Muse Image, Ideogram 4.0, P-Image-Ideogram, Qwen-Image API+local, Veo, Sora, LTX-2.5, Hailuo 2.3, MiniMax H3 / H3 Max, Fal H3 Max (+ Fal H3 Max Lip Sync, fal-h3-max-lipsync), Wan, Runway Gen-4.5, Nemotron Omni, Cosmos 3 Reasoner, Cosmos 3 Generator, Kimi K2.6/K2.7/K3, Qwen3.8 (+ Qwen3.8-Omni-Flash, qwen3.8-omni-flash), Hy4 preview (hy4-preview TokenHub + hy4-preview-local vLLM/SGLang), GLM-5.3-FlashX (glm-5.3-flashx, Zhipu/Z.ai — new tryon.api.zai), Ternary Bonsai 2 27B (ternary-bonsai-2-27b, PrismML local server — new tryon.models.ternary_bonsai), P-Video-2-Pro (p-video-2-pro, Pruna MiniMax H3-based), BEN2, dedicated cloud VTON (flux-vto, google-vton, outfitanyone-plus, photoroom-vton, photoroom-virtual-model, nova-canvas, kling-ai, FASHN, p-image-tryon, Segmind) plus composition try-on (nano-banana-2-lite, qwen-image, muse-image) and local dedicated VTON (leffa, catvton). Full table: CLI --help / registry.


Maintenance​

When adding a row: vendor, modality, Path A/B, proposed registry id, license/API URL, status.
When shipping: move to Shipped, delete the next row, note the registry id.
Re-survey NVIDIA: build.nvidia.com/models + Cosmos / Nemotron blogs. Do not paste the entire NIM catalog.
Re-survey VTON: Vertex Imagen try-on, DashScope OutfitAnyone, Photoroom/Pixelcut, Hugging Face CatVTON / Leffa / IDM-VTON. Prefer first-party APIs over aggregators.