Skip to main content

MCP Server

OpenTryOn ships a Model Context Protocol server under mcp-server/. Every model in tryon.cli.registry becomes an MCP tool automatically — the same surface as the opentryon CLI, via tryon.cli.runner.invoke_model().

Current release: OpenTryOn v0.0.5+ (pip install -U opentryon).

This page is the Docusaurus guide for the server. Keep it next to:

Why it matters​

  • Agents in Cursor, Claude Desktop, or TryOn Studio call try-on, generate, edit, video, understand, and bg-remove tools directly.
  • Studio chat goes through planner_agent: it classifies intent, then runs a filtered slice of those same registry tools via invoke_model. Capability screens skip the planner and call the model tools themselves.
  • New registry models appear as tools with zero hand-written MCP wrappers.
  • CLI and MCP cannot drift — one runner, one registry.

Install​

cd opentryon
pip install -e . # core (API-backed) models
# optional GPU extras (leffa, catvton, kimi-vl, qwen3.8, ben2, …):
pip install -e ".[local]"

cd mcp-server
pip install -r requirements.txt

Copy repo-root env.template to .env and fill the keys you plan to use. The server and every adapter read that same file.

Run​

# stdio — what Claude Desktop / Cursor expect
python server.py

# streamable-HTTP — required by TryOn Studio
python server.py --transport http --host 127.0.0.1 --port 8000

On startup the server prints a configuration report to stderr (which keys are set, which models are ready) — the same text opentryon_status returns at runtime.

TransportTypical clientEndpoint
stdio (default)Cursor, Claude Desktopprocess stdin/stdout
HTTPTryOn Studio, FastMCP Client over the networkhttp://127.0.0.1:8000/mcp

Studio’s only URL is OPENTRYON_MCP_URL=http://127.0.0.1:8000/mcp. Remote MCP hosts are out of scope for Studio.

Clients​

TryOn Studio​

HTTP MCP plus a Next.js UI (Agent, Connect, Image, VTON, Understand, Video, BG Remove). Full setup: TryOn Studio.

Cursor​

Add to .cursor/mcp.json (project) or ~/.cursor/mcp.json (global). Example: mcp-server/examples/cursor_mcp_config.json.

{
"mcpServers": {
"opentryon": {
"command": "python",
"args": ["/absolute/path/to/opentryon/mcp-server/server.py"]
}
}
}

Claude Desktop​

Add to claude_desktop_config.json. Example: mcp-server/examples/claude_desktop_config.json.

{
"mcpServers": {
"opentryon": {
"command": "python",
"args": ["/absolute/path/to/opentryon/mcp-server/server.py"]
}
}
}

Python (FastMCP client)​

See mcp-server/examples/example_usage.py.

import asyncio
from fastmcp import Client

async def main():
async with Client("server.py") as client:
result = await client.call_tool("vton_flux_vto", {
"person": "model.jpg",
"garment": "garment.jpg",
"dry_run": True,
})
print(result.data)

asyncio.run(main())

Discovery, keys, and the planner​

Always available, independent of which models you configured:

ToolRole
list_opentryon_toolsServices, models, MCP tool names, env readiness
opentryon_statusSame human-readable report printed on startup
list_api_keys / set_api_keysInspect or upsert host .env keys (never returns secret values). Studio Connect uses these
planner_agentStudio Agent chat entrypoint. Cheap LLM classifies intent, then invoke_model on a filtered slice. Planner Agent

Every generated model tool also accepts dry_run and output_dir, matching the CLI.

Selected tools​

The tables below highlight newer families. The complete generated list lives in mcp-server/README.md and grows with tryon/cli/registry.py.

Understand tools (including Qwen3.8 and Hy4)​

Multimodal image/video understanding tools include Kimi, LLaVA-NeXT, the Qwen3.8 dual path (+ Qwen3.8-Omni-Flash for audio), GLM-5.3-FlashX, Ternary Bonsai 2 27B (local server), and Hy4 preview (TokenHub LLM + local vLLM/SGLang):

MCP toolBackendNeeds
understand_qwen3_8_maxDashScope Qwen3.8-Max (text/image/video, thinking + reasoning_effort)DASHSCOPE_API_KEY
understand_qwen3_8_omni_flashDashScope Qwen3.8-Omni-Flash (adds audio; text/image/audio/video, 1M context)DASHSCOPE_API_KEY
understand_qwen3_8Local Qwen/Qwen3.8-27Bpip install opentryon[local] + GPU
understand_glm_5_3_flashxZ.ai GLM-5.3-FlashX (text/image/video, 200 tok/s)ZAI_API_KEY
understand_ternary_bonsai_2_27bPrismML Ternary Bonsai 2 27B (self-hosted llama.cpp/MLX server)BONSAI_BASE_URL (default localhost:8080)
understand_hy4_previewTencent Hy4 preview (TokenHub LLM)TOKENHUB_API_KEY
understand_hy4_preview_localHy4 via local vLLM/SGLang OpenAI serverHY4_BASE_URL (default localhost:8000)

Qwen-Image tools (generate / edit / VTON)​

Same DashScope key as Qwen3.8-Max. Image generation is Qwen-Image, not the Qwen3.8 VLM:

MCP toolBackendNeeds
generate_qwen_imageQwen-Image 3.0 T2I (default qwen-image-3.0-pro)DASHSCOPE_API_KEY
edit_qwen_imageQwen-Image I2I (1–3 refs)DASHSCOPE_API_KEY
vton_qwen_imagePerson + garment compositionDASHSCOPE_API_KEY

Typical loop: understand_qwen3_8_max captions a garment, then generate_qwen_image or vton_qwen_image uses that description.

Local Diffusers twin (pip install opentryon[local] + CUDA):

MCP toolBackendNeeds
generate_qwen_image_localQwen/Qwen-Image-2512 T2IGPU + recent Diffusers
edit_qwen_image_localQwen/Qwen-Image-Edit-2511 I2IGPU + recent Diffusers
vton_qwen_image_localEdit-Plus person + garmentGPU + recent Diffusers

See Qwen-Image local.

Local dedicated VTON (Leffa + CatVTON)​

MCP toolBackendNeeds
vton_leffafranciszzj/Leffa (CVPR 2025)GPU + opentryon[local]
vton_catvtonzhengchong/CatVTON (ICLR 2025, CC BY-NC-SA)GPU + opentryon[local]

See Leffa and CatVTON.

Qwen3.8 is a native multimodal / coding / agent family; OpenTryOn’s MCP tools expose the understand entry point (image and/or video + prompt). Full capability notes: Qwen3.8-Max, Qwen-Image, and Qwen3.8 local.

MiniMax H3 tools (video)​

Same MINIMAX_API_KEY as Hailuo 2.3. H3 is a dual path (hosted V2 vs local Diffusers); Hailuo 2.3 stays API-only. H3 Max is the hosted fast variant (no local twin).

MCP toolBackendNeeds
video_generate_minimax_h3MiniMax H3 official V2 API (T2V / I2V / R2V, 4–15s, 768P/2K)MINIMAX_API_KEY
video_generate_minimax_h3_maxMiniMax H3 Max (fast V2; T2V / I2V, 5–15s, 480P/768P)MINIMAX_API_KEY
video_generate_fal_h3_maxFal-hosted H3 Max (T2V / I2V / R2V, 5–15s, 480P/768P)FAL_KEY
video_generate_fal_h3_max_lipsyncFal-hosted H3 Max Lip Sync (image + audio → dubbed video, 480P–2K)FAL_KEY
video_generate_minimax_h3_localOpen-weight MiniMaxAI/MiniMax-H3 (768p H3-Base)pip install opentryon[local] + CUDA + Diffusers from main
video_generate_hailuo_2_3MiniMax Hailuo 2.3 (V1)MINIMAX_API_KEY

See MiniMax H3 API, MiniMax H3 Max (Fal), and MiniMax H3 local.

NVIDIA NIM tools (understand + video)​

Same NVIDIA_API_KEY for Nemotron Omni, Cosmos 3 Reasoner, and Cosmos 3 Generator.

MCP toolBackendNeeds
understand_nemotron_omniNemotron 3 Nano Omni (image / video / audio)NVIDIA_API_KEY
understand_cosmos3_reasonerCosmos 3 Reasoner (physical-world VLM)NVIDIA_API_KEY
video_generate_cosmos3Cosmos 3 Generator nano (T2V / I2V)NVIDIA_API_KEY

See NVIDIA NIM.

Google Virtual Try-On (Vertex)​

Dedicated person + product try-on. Not GEMINI_API_KEY / Nano Banana.

MCP toolBackendNeeds
vton_google_vtonVertex virtual-try-on-001GOOGLE_CLOUD_PROJECT + ADC

See Google Virtual Try-On.

OutfitAnyone-Plus (DashScope, Beijing)​

Dedicated Alibaba try-on. Not Qwen-Image composition. Needs a China Beijing-region DASHSCOPE_API_KEY.

MCP toolBackendNeeds
vton_outfitanyone_plusaitryon-plusBeijing DASHSCOPE_API_KEY

See OutfitAnyone-Plus.

Photoroom (Virtual Try-On / Virtual Model)​

Image Editing API Plus. Same PHOTOROOM_API_KEY. Prefix with sandbox_ for watermarked tests.

MCP toolBackendNeeds
vton_photoroom_vtonShopper try-on (virtualModel.model.custom)PHOTOROOM_API_KEY
vton_photoroom_virtual_modelCatalog on-model (preset or custom)PHOTOROOM_API_KEY

See Photoroom.

Muse Image tools (generate / edit / VTON)​

First-party Meta Model API (MODEL_API_KEY). No local twin. Muse Video is not on the API yet.

MCP toolBackendNeeds
generate_muse_imageMuse Image T2I (muse-image-1.0)MODEL_API_KEY
edit_muse_imageMuse Image I2I / multi-refMODEL_API_KEY
vton_muse_imagePerson + garment compositionMODEL_API_KEY

See Muse Image and Muse Video (not available).

P-Image-Ideogram (generate)​

Same PRUNA_API_KEY as the rest of the Pruna family. Not Ideogram 4.0 (generate_ideogram / IDEOGRAM_API_KEY).

MCP toolBackendNeeds
generate_p_image_ideogramPruna Model: p-image-ideogram (thinking very-low–very-high, 1K/2K)PRUNA_API_KEY

See P-Image-Ideogram.