MiniMax H3 Max (Fal)
OpenTryOn’s first third-party hoster adapter. Fal jointly released MiniMax H3 Max with MiniMax and hosts a post-trained stack with T2V, I2V, reference-to-video, and a dedicated lip-sync/dubbing endpoint.
This is not the MiniMax first-party V2 API. Official MiniMax H3 Max (--model minimax-h3-max, MINIMAX_API_KEY) is T2V / first-last I2V only. Use Fal when you want R2V, lip-sync, or Fal’s queue.
| CLI model | Host | Modes | Key |
|---|---|---|---|
minimax-h3-max | MiniMax V2 | T2V, I2V | MINIMAX_API_KEY |
fal-h3-max | Fal queue | T2V, I2V, R2V | FAL_KEY |
fal-h3-max-lipsync | Fal queue | Image + audio → lip-synced video | FAL_KEY |
minimax-h3 | MiniMax V2 | T2V, I2V, R2V (Python), 2K | MINIMAX_API_KEY |
No open weights for H3 Max. Local H3-Base is --model minimax-h3-local.
Docs: T2V, I2V, R2V, Lip Sync, queue.
Environment
export FAL_KEY=... # https://fal.ai/dashboard/keys
# export FAL_API_KEY=... # alias
# export FAL_QUEUE_BASE_URL=https://queue.fal.run
Auth header is Authorization: Key {FAL_KEY} (not Bearer).
CLI
opentryon video-generate --model fal-h3-max \
--prompt "A fashion model walking a runway at dusk, camera tracking" \
--duration 5 --resolution 768P --ratio 16:9
opentryon video-generate --model fal-h3-max \
--image look.jpg --prompt "Gentle fabric motion as the model turns" \
--duration 6
opentryon video-generate --model fal-h3-max \
--prompt "Walk toward camera, stop on the mark" \
--last-frame end.jpg
opentryon video-generate --model fal-h3-max \
--prompt "Image 1 is the model. Keep her identity while she walks the runway." \
--reference-image look.jpg \
--duration 5 --resolution 768P
# Lip-sync / dubbing: portrait + audio in, matched mouth movement out (no prompt)
opentryon video-generate --model fal-h3-max-lipsync \
--image portrait.jpg --audio line.wav \
--resolution 1080P --enable-transcription
| Flag | Notes |
|---|---|
--duration | Integer 5–15 seconds (default 5) |
--resolution | 480P or 768P (default 768P) |
--ratio | T2V: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 (not adaptive). R2V may use adaptive. I2V follows the keyframe. |
--image | First frame (switches to I2V) |
--last-frame | Last frame (alone or with --image) |
--reference-image / --reference-video / --reference-audio | R2V lists. At most 12 files total. Audio cannot be the only reference. Mutually exclusive with --image. |
--prompt-expansion | balanced (default) or quality |
--no-safety-checker | Disable Fal’s safety checker |
--seed | Optional |
fal-h3-max-lipsync flags
No --prompt — only an image and an audio track. Resolution ceiling is higher than T2V/I2V/R2V (up to 2K).
| Flag | Notes |
|---|---|
--image | Portrait to animate (aspect ratio must be 0.4–2.5) — required |
--audio | Track to sync mouth movement to (≥5s; clipped to ~14.8s) — required |
--resolution | 480P, 768P, 1080P, or 2K (default 768P) |
--enable-transcription | Transcribe the audio to guide lip-sync accuracy |
--no-safety-checker | Disable Fal’s safety checker |
--seed | Optional |
Python
from tryon.api.fal import FalH3MaxAdapter
adapter = FalH3MaxAdapter()
open("t2v.mp4", "wb").write(
adapter.generate_text_to_video(
prompt="A fashion model walking a runway at dusk",
duration=5,
resolution="768P",
ratio="16:9",
)
)
open("i2v.mp4", "wb").write(
adapter.generate_image_to_video(
image="look.jpg",
prompt="Gentle turn toward the camera",
last_frame="end.jpg",
)
)
open("r2v.mp4", "wb").write(
adapter.generate_text_to_video(
prompt="Image 1 is the model. Keep her identity while she walks.",
reference_image=["look.jpg"],
)
)
open("lipsync.mp4", "wb").write(
adapter.generate_lip_sync(
image="portrait.jpg",
audio="line.wav",
resolution="1080P",
)
)
Notes
- Jobs go through Fal’s queue (
POST https://queue.fal.run/minimax/h3-max/...→ pollstatus_url→ downloadvideo.url). - MCP tools:
video_generate_fal_h3_maxandvideo_generate_fal_h3_max_lipsync(generated from the registry). - Planner:
fal h3 max/fal-h3-maxpin T2V/I2V/R2V.h3 max lip sync/fal-h3-max-lipsyncpin the dubbing endpoint. A bareh3 maxstill pins first-partyminimax-h3-max.