window.WORKFLOW_DATA = { "prompts": { "fl2va": "You are an expert prompt writer for MiniMax-H3 FL2VA / First-and-Last-Frame-to-Video generation.\n\nConvert the user's input directly into a precise, generation-oriented FL2VA prompt.\n\nOutput ONLY:\n\nFL2VA:\n\nsubject_definitions:\n...\n\nsummary:\n...\n\nretention_analysis:\n...\n\ndetailed_description:\n...\n\noverall_soundscape:\n...\n\nnon_diegetic_music:\n...\n\nWrite structural prose in English.\n\n## CORE RULES\n\nDescribe ONLY the target video.\n\nUse only:\n\n1. Explicit user instructions.\n2. Relevant visual information from assigned First/Last Frame references.\n3. Explicit motion/camera/timing from assigned Video references.\n4. Explicit sound/music from assigned Audio references.\n\nFirst/Last Frame references define the required visual state and appearance at the beginning and end of the video. They do NOT define intermediate actions, motion, camera movement, timing, dialogue, sound, or music unless explicitly assigned as such.\n\nNever invent events, reactions, gestures, sounds, camera movements, dialogue, transitions, or endings.\n\nThe user's timeline is authoritative for WHAT happens.\n\n## FIRST / LAST FRAME RULE\n\nPreserve the visual identity and relevant appearance of the assigned frames.\n\nThe first frame represents the starting visual state.\n\nThe last frame represents the ending visual state.\n\nDo not describe the first or last frame as an action unless the user explicitly specifies that action.\n\nDo not invent intermediate motion solely to explain how the first frame becomes the last frame.\n\nUse only the user's explicit timeline for intermediate actions and motion.\n\n## TIMELINE\n\nPreserve every explicit timestamp exactly.\n\nPreserve every explicit:\n\n* action\n* motion\n* camera instruction\n* speech\n* narration\n* voiceover\n* lyric\n* visible text\n* sound\n\nUse `[Shot 1]` unless multiple shots are explicitly specified.\n\nKeep each interval concise. Do not merge away explicit events.\n\n## SPEECH — CRITICAL\n\nA speech event exists ONLY when:\n\n1. The surrounding text explicitly indicates speech, narration, voiceover, or singing.\n2. The actual words are enclosed in quotation marks.\n\nAccepted quotation marks:\n\n「...」\n“...”\n\"...\"\n‘...’\n'...'\n\nEvery valid quoted speech event is LOCKED CONTENT.\n\nFor every valid speech event:\n\n* preserve exact wording\n* preserve exact punctuation\n* preserve original language\n* preserve speaker\n* preserve exact timestamp\n* output it exactly once\n* place it in the corresponding timeline interval\n* wrap the actual spoken words in `...`\n\nExample:\n\n`[10-15s] She says in Japanese 「Example phrase.」`\n\n→\n\n`[10-15s] She says in Japanese: Example phrase.`\n\nThe example is generic and must never be inserted unless supplied by the user.\n\nNEVER:\n\n* translate\n* romanize\n* paraphrase\n* summarize\n* shorten\n* expand\n* correct\n* replace\n* omit\n* duplicate\n\nDo NOT generate dialogue from statements such as:\n\n* She speaks Japanese.\n* She talks to the camera.\n* She says something.\n* Japanese dialogue occurs.\n\nWithout quoted words, there is NO dialogue.\n\nIf an interval has no valid speech, write exactly:\n\n`No dialogue or narration.`\n\nIf an interval contains valid speech, do NOT write that phrase.\n\n## VISIBLE TEXT\n\nQuoted text is NOT automatically speech.\n\nIf the user identifies it as a subtitle, sign, label, title, poster, screen text, written message, or other visible text, preserve it exactly in the original language and do not speak it.\n\n## REFERENCE RULES\n\n### subject_definitions\n\nUse only user-provided subject information and relevant visual information from the assigned First/Last Frame references.\n\n### summary\n\nBriefly summarize the target video. Do not replace or omit explicit dialogue.\n\n### retention_analysis\n\nMention only relevant visual information retained from the assigned First/Last Frame references. Do not add unsupported actions, motion, camera behavior, sound, or dialogue.\n\n### detailed_description\n\nThis is the authoritative timeline.\n\nEvery explicit user event MUST appear here.\n\nEvery valid speech event MUST appear here exactly once with its original wording and timestamp.\n\nDescribe intermediate motion only when explicitly provided by the user or an assigned Video reference.\n\n### overall_soundscape\n\nUse only explicitly specified or explicitly referenced sound.\n\nIf none:\n\n`N/A`\n\n### non_diegetic_music\n\nUse only explicitly specified or explicitly referenced music.\n\nIf none:\n\n`N/A`\n\n## CAMERA / MOTION / TIMING / SOUND\n\nCamera behavior:\nONLY from user instructions or assigned Video references.\n\nActions and motion:\nONLY from user instructions or assigned Video references.\n\nTiming:\nONLY from user timestamps or assigned Video references.\n\nSound:\nONLY from user instructions or assigned Audio references.\n\nMusic:\nONLY from user instructions or assigned Audio references.\n\nDo not invent camera movement or motion to connect the First and Last Frames.\n\n## ENDING\n\nThe last frame defines the required final visual state when a Last Frame reference is provided.\n\nDo not invent any additional ending action, transition, fade, freeze, final hold, credits, or reaction beyond the user's instructions.\n\n## FINAL CHECK\n\nSilently verify:\n\n* Every explicit timeline interval is present.\n* Every explicit action/motion/camera instruction is present.\n* The First Frame visual state is preserved.\n* The Last Frame visual state is preserved.\n* No unsupported intermediate motion was invented from the frames.\n* Every valid quoted speech event is present.\n* Every speech event appears exactly once.\n* Every speech event retains exact wording, punctuation, language, speaker, and timestamp.\n* `` contains only the actual spoken words.\n* No unquoted text became dialogue.\n* No dialogue was invented.\n* No unsupported sound, music, camera behavior, motion, reaction, or transition was invented.\n* Output contains ONLY the required FL2VA structure.\n\nNever output reasoning, analysis, warnings, explanations, or commentary.\nNever output text before `FL2VA:` or after the completed prompt.\n", "ref2va": "You are an expert prompt writer for MiniMax-H3 Ref2VA / Full-Reference video generation.\n\nConvert the user's input directly into a precise generation-oriented Ref2VA prompt.\n\nOutput ONLY:\n\nRef2VA:\n\nsubject_definitions:\n...\n\nsummary:\n...\n\nretention_analysis:\n...\n\ndetailed_description:\n...\n\noverall_soundscape:\n...\n\nnon_diegetic_music:\n...\n\nWrite structural prose in English.\n\n## CORE RULES\n\nDescribe ONLY the target video.\n\nUse only:\n\n1. Explicit user instructions.\n2. Relevant visual information from assigned Picture references.\n3. Explicit motion/camera/timing from assigned Video references.\n4. Explicit sound/music from assigned Audio references.\n\nPictures provide visual information only. Never derive actions, motion, camera behavior, timing, dialogue, sound, or music from Pictures.\n\nNever invent events, reactions, gestures, sounds, camera movements, dialogue, transitions, or endings.\n\nThe user's timeline is authoritative for WHAT happens.\n\n## TIMELINE\n\nPreserve every explicit timestamp exactly.\n\nPreserve every explicit:\n\n* action\n* motion\n* camera instruction\n* speech\n* narration\n* voiceover\n* lyric\n* visible text\n* sound\n\nUse `[Shot 1]` unless multiple shots are explicitly specified.\n\nKeep each interval concise. Do not merge away explicit events.\n\n## SPEECH — CRITICAL\n\nA speech event exists ONLY when:\n\n1. The surrounding text explicitly indicates speech, narration, voiceover, or singing.\n2. The actual words are enclosed in quotation marks.\n\nAccepted quotation marks:\n\n「...」\n“...”\n\"...\"\n‘...’\n'...'\n\nEvery valid quoted speech event is LOCKED CONTENT.\n\nFor every valid speech event:\n\n* preserve exact wording\n* preserve exact punctuation\n* preserve original language\n* preserve speaker\n* preserve exact timestamp\n* output it exactly once\n* place it in the corresponding timeline interval\n* wrap the actual spoken words in `...`\n\nExample:\n\n`[10-15s] She says in Japanese 「Example phrase.」`\n\n→\n\n`[10-15s] She says in Japanese: Example phrase.`\n\nThe example is generic and must never be inserted unless supplied by the user.\n\nNEVER:\n\n* translate\n* romanize\n* paraphrase\n* summarize\n* shorten\n* expand\n* correct\n* replace\n* omit\n* duplicate\n\nDo NOT generate dialogue from statements such as:\n\n* She speaks Japanese.\n* She talks to the camera.\n* She says something.\n* Japanese dialogue occurs.\n\nWithout quoted words, there is NO dialogue.\n\nIf an interval has no valid speech, write exactly:\n\n`No dialogue or narration.`\n\nIf an interval contains valid speech, do NOT write that phrase.\n\n## VISIBLE TEXT\n\nQuoted text is NOT automatically speech.\n\nIf the user identifies it as a subtitle, sign, label, title, poster, screen text, written message, or other visible text, preserve it exactly in the original language and do not speak it.\n\n## REFERENCE RULES\n\n### subject_definitions\n\nUse only user-provided subject information and relevant visual information from Pictures.\n\n### summary\n\nBriefly summarize the target video. Do not replace or omit explicit dialogue.\n\n### retention_analysis\n\nMention only relevant visual information retained from Pictures. Do not add actions, motion, camera behavior, sound, or dialogue.\n\n### detailed_description\n\nThis is the authoritative timeline.\n\nEvery explicit user event MUST appear here.\n\nEvery valid speech event MUST appear here exactly once with its original wording and timestamp.\n\n### overall_soundscape\n\nUse only explicitly specified or explicitly referenced sound.\n\nIf none:\n\n`N/A`\n\n### non_diegetic_music\n\nUse only explicitly specified or explicitly referenced music.\n\nIf none:\n\n`N/A`\n\n## CAMERA / MOTION / TIMING / SOUND\n\nCamera behavior:\nONLY from user instructions or assigned Video references.\n\nActions and motion:\nONLY from user instructions or assigned Video references.\n\nTiming:\nONLY from user timestamps or assigned Video references.\n\nSound:\nONLY from user instructions or assigned Audio references.\n\nMusic:\nONLY from user instructions or assigned Audio references.\n\n## ENDING\n\nDo not invent an ending, transition, fade, freeze, final hold, credits, or additional reaction unless explicitly specified.\n\n## FINAL CHECK\n\nSilently verify:\n\n* Every explicit timeline interval is present.\n* Every explicit action/motion/camera instruction is present.\n* Every valid quoted speech event is present.\n* Every speech event appears exactly once.\n* Every speech event retains exact wording, punctuation, language, speaker, and timestamp.\n* `` contains only the actual spoken words.\n* No unquoted text became dialogue.\n* No dialogue was invented.\n* No Picture reference created an event.\n* No unsupported sound, music, camera behavior, motion, reaction, or ending was invented.\n* Output contains ONLY the required Ref2VA structure.\n\nNever output reasoning, analysis, warnings, explanations, or commentary.\nNever output text before `Ref2VA:` or after the completed prompt.\n" }, "workflows": [ { "file": "workflows/MiniMax_int8_I2V-javano2609.1.2.json", "title": "Image-to-Video (FLF)", "subtitle": "First-Last-Frame generation with Extend, ControlNet-Union V2V and Latent Upscaler", "description": "Condition the video on the first AND last frames. Built on the INT8 FL2VA model with native Block Sparse Attention, Spectrum sampling, Lightx2v/Turbo LoRAs, native Extend, Fun ControlNet Union (OpenPose/Canny/Depth) and low-res -> latent upscale -> high-res decode.", "features": [ "First + Last Frame", "INT8", "Block Sparse Attention", "Fun ControlNet Union", "Native Extend", "Latent Upscaler", "Spectrum", "Lightx2v LoRA", "Group Bypass" ], "topLevelNodes": 68, "totalNodes": 274, "inventory": [ { "type": "GetNode", "count": 35 }, { "type": "ComfySwitchNode", "count": 29 }, { "type": "SetNode", "count": 19 }, { "type": "ComfyMathExpression", "count": 13 }, { "type": "ImpactIfNone", "count": 8 }, { "type": "PrimitiveStringMultiline", "count": 6 }, { "type": "easy int", "count": 6 }, { "type": "Reroute", "count": 6 }, { "type": "LoraLoaderModelOnly", "count": 6 }, { "type": "Note", "count": 5 }, { "type": "VHS_VideoCombine", "count": 5 }, { "type": "LoadImage", "count": 5 }, { "type": "easy anythingIndexSwitch", "count": 5 }, { "type": "GetImageRangeFromBatch", "count": 5 }, { "type": "PrimitiveBoolean", "count": 4 }, { "type": "RandomNoise", "count": 4 }, { "type": "BasicScheduler", "count": 4 }, { "type": "LTXVConcatAVLatent", "count": 4 }, { "type": "LTXVSeparateAVLatent", "count": 4 }, { "type": "MiniMaxH3ImageToVideo", "count": 4 }, { "type": "MMH3SplitUpscale", "count": 4 }, { "type": "MiniMaxH3FunControlNetApply", "count": 4 }, { "type": "TrimAudioDuration", "count": 4 }, { "type": "Group Controller", "count": 3 }, { "type": "ImageResizeKJv2", "count": 3 }, { "type": "PreviewAny", "count": 2 }, { "type": "KSamplerSelect", "count": 2 }, { "type": "AudioEnhancementNode", "count": 2 }, { "type": "BasicGuider", "count": 2 }, { "type": "VAEDecodeAudio", "count": 2 }, { "type": "RTXVideoSuperResolution", "count": 2 }, { "type": "MinimaxH3LatentUpscaler3D", "count": 2 }, { "type": "VAEDecode", "count": 2 }, { "type": "SamplerCustomAdvanced", "count": 2 }, { "type": "MMH3TemporalSplitParamsV10", "count": 2 }, { "type": "MMH3SpatialSplitParamsV10", "count": 2 }, { "type": "ManualSigmas", "count": 2 }, { "type": "LTXVAudioVAEEncode", "count": 2 }, { "type": "SolidMask", "count": 2 }, { "type": "SetLatentNoiseMask", "count": 2 }, { "type": "MiniMaxH3AddGuide", "count": 2 }, { "type": "GetImageSizeAndCount", "count": 2 }, { "type": "OllamaGenerateV2", "count": 2 }, { "type": "ImpactConditionalBranchSelMode", "count": 2 }, { "type": "VAELoader", "count": 2 }, { "type": "MarkdownNote", "count": 1 }, { "type": "35348f87-ce77-463d-914a-9b04b1a6e09b", "count": 1 }, { "type": "2f377a81-7be4-4d8b-bba5-464d1e5329b1", "count": 1 }, { "type": "50874503-87e5-4f15-a304-7c702a3cb78d", "count": 1 }, { "type": "a6811221-294a-4519-a645-6e0b07df229e", "count": 1 }, { "type": "cc4f3c17-635a-45cc-bf44-6d1c3af27c3b", "count": 1 }, { "type": "ResolutionSelector", "count": 1 }, { "type": "LoadVideoUI", "count": 1 }, { "type": "ModelPreviewOverrideKJ", "count": 1 }, { "type": "850d0cb6-8c18-4748-a2ff-97da873fffb0", "count": 1 }, { "type": "e69c9dec-17d5-4d89-8bf3-1a480552d301", "count": 1 }, { "type": "7918e590-34c0-4440-b5a3-2cd2b4e767fa", "count": 1 }, { "type": "4439dde7-fcea-4941-8996-01e96a0e87a1", "count": 1 }, { "type": "VHS_LoadVideo", "count": 1 }, { "type": "fdcfce6d-b7d6-45cb-a01c-59db5101e70a", "count": 1 }, { "type": "easy forLoopEnd", "count": 1 }, { "type": "AudioConcat", "count": 1 }, { "type": "easy forLoopStart", "count": 1 }, { "type": "ImageBatchExtendWithOverlap", "count": 1 }, { "type": "easy boolean", "count": 1 }, { "type": "OllamaConnectivityV2", "count": 1 }, { "type": "OllamaOptionsV2", "count": 1 }, { "type": "easy convertAnything", "count": 1 }, { "type": "CustomCombo", "count": 1 }, { "type": "CLIPLoader", "count": 1 }, { "type": "MiniMaxH3MemoryEfficientSageAttentionPatch", "count": 1 }, { "type": "MiniMaxH3SigmaShift", "count": 1 }, { "type": "SpectrumApplyMiniMaxH3", "count": 1 }, { "type": "UNETLoader", "count": 1 }, { "type": "PathchSageAttentionKJ", "count": 1 }, { "type": "ModelPatchTorchSettings", "count": 1 }, { "type": "ModelAttentionBackend", "count": 1 }, { "type": "ModelPatchLoader", "count": 1 }, { "type": "BlockSparseAttention", "count": 1 }, { "type": "RMBG", "count": 1 }, { "type": "DWPreprocessor", "count": 1 }, { "type": "ResizeImageMaskNode", "count": 1 }, { "type": "CannyEdgePreprocessor", "count": 1 }, { "type": "DepthAnythingPreprocessor", "count": 1 } ], "models": [ "H3\\minimax_h3_fl2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors", "minimax_h3_audio_vae_fp32.safetensors", "minimax_h3_fl2va_pruned_int8_convrot.safetensors", "minimax_h3_fun_controlnet_union_pruned_int8_convrot.safetensors", "minimax_h3_latent_upscaler_3d_fp16.safetensors", "minimax_h3_video_vae_int8_convrot.safetensors", "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors", "taeh3.safetensors" ], "notes": [ "| megapixels | Aspect | Output (multiple=32) | |---|---|---| | 0.2 | 16:9 | 608 x 352 | | 0.3 | 16:9 | 736 x 416 | | 0.4 | 16:9 | 864 x 480 | | 0.5 | 16:9 | 960 x 544 | | 0.6 | 16:9 | 1056 x 608 | | 0.7 | 16:9 | 1152 x 640 | | 0.8 | 16:9 | 1216 x 672 | | 0.9 | 16:9 | 1280 x 736 | | 0.98 | 16:9 | 1344 x 768 | | 1.0 | 16:9 | 1376 x 768 | | 1.2 | 16:9 | 1504 x 832 | | 1.5 | 16:9 | 1664 x 928 | | 1.8 | 16:9 | 1824 x 1024 | | 2.0 | 16:9 | 1920 x 1088 |", "Lightx2v Turbo-LoRA 8 step ver : shift_video 12.0 8 step 768p ver : shift_video 6.0 4 step ver : shift_video 6.0", "Comfyui_Minimax_h3_latent_Upscaler https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler Controlnet Union https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/controlnet Lightx2v Turbo-LoRA https://huggingface.co/lightx2v/Minimax-h3-Turbo Turbo-LoRA https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI PDD ACC LoRA https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras Spectrum https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3.git", "FL2VA does not support audio reference, so the audio from the Control Video will be ignored. Due to \"Start index is out of range\" error, only when using Controlnet for V2V during Extend, the generated video must be 22 frames (the offset) shorter than the Control Video.", "0: Depth 1: Canny 2: Openpose" ] }, { "file": "workflows/MiniMax_int8_R2V-javano2609.2.2.json", "title": "Reference-to-Video (R2V)", "subtitle": "Reference-guided generation with Ref LoRA, ControlNet-Union V2V, Extend and Upscaler", "description": "Animate from one or more reference images using the R2V Reference LoRA, with Fun ControlNet Union video control, native Extend, audio-VAE aware outputs and the Latent Upscaler. The most complete workflow in the collection.", "features": [ "Reference LoRA", "INT8", "Block Sparse Attention", "Fun ControlNet Union", "Native Extend", "Latent Upscaler", "Spectrum", "Lightx2v LoRA", "Group Bypass" ], "topLevelNodes": 86, "totalNodes": 370, "inventory": [ { "type": "GetNode", "count": 53 }, { "type": "ComfySwitchNode", "count": 40 }, { "type": "SetNode", "count": 28 }, { "type": "ComfyMathExpression", "count": 18 }, { "type": "GetImageRangeFromBatch", "count": 13 }, { "type": "Reroute", "count": 9 }, { "type": "ImpactIfNone", "count": 8 }, { "type": "LoraLoaderModelOnly", "count": 7 }, { "type": "ResizeImageMaskNode", "count": 7 }, { "type": "LoadImage", "count": 6 }, { "type": "VHS_VideoCombine", "count": 6 }, { "type": "PrimitiveStringMultiline", "count": 6 }, { "type": "LTXVSeparateAVLatent", "count": 6 }, { "type": "SetLatentNoiseMask", "count": 6 }, { "type": "MiniMaxH3ReferenceToVideo", "count": 6 }, { "type": "TrimAudioDuration", "count": 6 }, { "type": "Note", "count": 5 }, { "type": "easy int", "count": 5 }, { "type": "PrimitiveBoolean", "count": 5 }, { "type": "easy anythingIndexSwitch", "count": 4 }, { "type": "RandomNoise", "count": 4 }, { "type": "BasicScheduler", "count": 4 }, { "type": "VAEDecode", "count": 4 }, { "type": "LTXVConcatAVLatent", "count": 4 }, { "type": "MMH3SplitUpscale", "count": 4 }, { "type": "MiniMaxH3FunControlNetApply", "count": 4 }, { "type": "Group Controller", "count": 3 }, { "type": "ImageResizeKJv2", "count": 3 }, { "type": "CustomCombo", "count": 3 }, { "type": "PreviewAny", "count": 2 }, { "type": "LoadVideoUI", "count": 2 }, { "type": "ImageFromBatch", "count": 2 }, { "type": "OllamaGenerateV2", "count": 2 }, { "type": "ImpactConditionalBranchSelMode", "count": 2 }, { "type": "VAELoader", "count": 2 }, { "type": "SolidMask", "count": 2 }, { "type": "VAEDecodeAudio", "count": 2 }, { "type": "AudioEnhancementNode", "count": 2 }, { "type": "KSamplerSelect", "count": 2 }, { "type": "LTXVAudioVAEEncode", "count": 2 }, { "type": "BasicGuider", "count": 2 }, { "type": "RTXVideoSuperResolution", "count": 2 }, { "type": "SamplerCustomAdvanced", "count": 2 }, { "type": "MinimaxH3LatentUpscaler3D", "count": 2 }, { "type": "MMH3TemporalSplitParamsV10", "count": 2 }, { "type": "MMH3SpatialSplitParamsV10", "count": 2 }, { "type": "ManualSigmas", "count": 2 }, { "type": "VAEEncode", "count": 2 }, { "type": "GetLatentSizeAndCount", "count": 2 }, { "type": "GetImageSizeAndCount", "count": 2 }, { "type": "MiniMaxH3AddGuide", "count": 2 }, { "type": "MarkdownNote", "count": 1 }, { "type": "07119851-5428-48a3-a461-1ab3b068799c", "count": 1 }, { "type": "ModelPreviewOverrideKJ", "count": 1 }, { "type": "3a08944f-220c-4078-8681-63c0f260e548", "count": 1 }, { "type": "cc4f3c17-635a-45cc-bf44-6d1c3af27c3b", "count": 1 }, { "type": "3a663545-426e-47c6-9401-c96b59b6dc1a", "count": 1 }, { "type": "ResolutionSelector", "count": 1 }, { "type": "LoadAudioUI", "count": 1 }, { "type": "a04e3fae-482c-4f75-a675-282716b9c592", "count": 1 }, { "type": "d0a77711-a824-4803-a311-f80a6e561561", "count": 1 }, { "type": "72e01e94-1f68-42a8-bb51-b9fad6de26c5", "count": 1 }, { "type": "b574ad05-3d50-4304-9b1d-c2c708ad36cb", "count": 1 }, { "type": "82aa8b12-c7a8-422b-ba78-0f6e430e13e9", "count": 1 }, { "type": "VHS_LoadVideo", "count": 1 }, { "type": "60493b99-ebf9-4301-ae6d-e9bd428b5fbb", "count": 1 }, { "type": "SaveImageAdvanced", "count": 1 }, { "type": "BatchImagesNode", "count": 1 }, { "type": "86ff6a37-c563-40d5-8f86-7f95e77d4b23", "count": 1 }, { "type": "923cd7c7-0083-46a2-9e46-d1f1817b3ad8", "count": 1 }, { "type": "easy boolean", "count": 1 }, { "type": "OllamaOptionsV2", "count": 1 }, { "type": "OllamaConnectivityV2", "count": 1 }, { "type": "CLIPLoader", "count": 1 }, { "type": "MiniMaxH3MemoryEfficientSageAttentionPatch", "count": 1 }, { "type": "MiniMaxH3SigmaShift", "count": 1 }, { "type": "SpectrumApplyMiniMaxH3", "count": 1 }, { "type": "UNETLoader", "count": 1 }, { "type": "PathchSageAttentionKJ", "count": 1 }, { "type": "ModelPatchTorchSettings", "count": 1 }, { "type": "ModelAttentionBackend", "count": 1 }, { "type": "ModelPatchLoader", "count": 1 }, { "type": "BlockSparseAttention", "count": 1 }, { "type": "easy forLoopEnd", "count": 1 }, { "type": "AudioConcat", "count": 1 }, { "type": "ImageBatchExtendWithOverlap", "count": 1 }, { "type": "easy forLoopStart", "count": 1 }, { "type": "ColorMatchV2", "count": 1 }, { "type": "easy convertAnything", "count": 1 }, { "type": "RMBG", "count": 1 }, { "type": "DWPreprocessor", "count": 1 }, { "type": "CannyEdgePreprocessor", "count": 1 }, { "type": "DepthAnythingPreprocessor", "count": 1 }, { "type": "CLIPTextEncode", "count": 1 }, { "type": "CheckpointLoaderSimple", "count": 1 }, { "type": "SAM3_TrackToMask", "count": 1 }, { "type": "SAM3_VideoTrack", "count": 1 }, { "type": "GrowMaskWithBlur", "count": 1 }, { "type": "DrawMaskOnImage", "count": 1 }, { "type": "ImpactSwitch", "count": 1 } ], "models": [ "H3\\minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors", "H3\\minimax_h3_ref_lora_rank_256_bf16.safetensors", "minimax_h3_audio_vae_fp32.safetensors", "minimax_h3_fl2va_pruned_int8_convrot.safetensors", "minimax_h3_fun_controlnet_union_pruned_int8_convrot.safetensors", "minimax_h3_latent_upscaler_3d_fp16.safetensors", "minimax_h3_video_vae_int8_convrot.safetensors", "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors", "sam3.1_multiplex_fp16.safetensors", "taeh3.safetensors" ], "notes": [ "Lightx2v Turbo-LoRA 8 step ver : shift_video 12.0 4 step ver : shift_video 6.0", "Ref LoRA https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras # When using Ref LoRA, inference can be performed using the Fl2VA model. Controlnet Union https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/controlnet Comfyui_Minimax_h3_latent_Upscaler https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler Lightx2v Turbo-LoRA https://huggingface.co/lightx2v/Minimax-h3-Turbo Turbo-LoRA https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI PDD ACC LoRA https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras Spectrum-MiniMax-H3 https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3.git", "| megapixels | Aspect | Output (multiple=32) | |---|---|---| | 0.2 | 16:9 | 608 x 352 | | 0.3 | 16:9 | 736 x 416 | | 0.4 | 16:9 | 864 x 480 | | 0.5 | 16:9 | 960 x 544 | | 0.6 | 16:9 | 1056 x 608 | | 0.7 | 16:9 | 1152 x 640 | | 0.8 | 16:9 | 1216 x 672 | | 0.9 | 16:9 | 1280 x 736 | | 0.98 | 16:9 | 1344 x 768 | | 1.0 | 16:9 | 1376 x 768 | | 1.2 | 16:9 | 1504 x 832 | | 1.5 | 16:9 | 1664 x 928 | | 1.8 | 16:9 | 1824 x 1024 | | 2.0 | 16:9 | 1920 x 1088 |", "Due to \"Start index is out of range\" error, only when using Controlnet or Edit for V2V during Extend, the generated video must be 22 frames (the offset) shorter than the Control Video or Target Video.", "If an \"aimdo memory compile error\" occurs, please launch ComfyUI with the \"--disable-comfy-compiler\" flag." ] }, { "file": "workflows/MiniMax_int8_Bridge-javano2609.1.1.json", "title": "Bridge", "subtitle": "Generate a transition clip between Video A and Video B", "description": "Takes two clips and synthesizes a bridge that connects them: Video A -> generated cut -> Video B. Uses the R2V pipeline with Ref LoRA and keeps all clips at the same size.", "features": [ "Video A -> Bridge -> Video B", "R2V pipeline", "Ref LoRA", "INT8", "Latent Upscaler" ], "topLevelNodes": 29, "totalNodes": 128, "inventory": [ { "type": "ComfySwitchNode", "count": 18 }, { "type": "SetNode", "count": 12 }, { "type": "GetNode", "count": 12 }, { "type": "ComfyMathExpression", "count": 7 }, { "type": "LoraLoaderModelOnly", "count": 7 }, { "type": "LoadImage", "count": 4 }, { "type": "MiniMaxH3AddGuide", "count": 4 }, { "type": "ImageFromBatch", "count": 4 }, { "type": "Note", "count": 3 }, { "type": "AudioConcat", "count": 3 }, { "type": "VHS_VideoCombine", "count": 2 }, { "type": "LoadVideoUI", "count": 2 }, { "type": "MiniMaxH3ReferenceToVideo", "count": 2 }, { "type": "ImageResizeKJv2", "count": 2 }, { "type": "BasicScheduler", "count": 2 }, { "type": "TrimAudioDuration", "count": 2 }, { "type": "BatchImagesNode", "count": 2 }, { "type": "MMH3SplitUpscale", "count": 2 }, { "type": "ResizeImageMaskNode", "count": 2 }, { "type": "VAELoader", "count": 2 }, { "type": "8f69a255-ac28-4d17-9ddd-4e2f270b2235", "count": 1 }, { "type": "PreviewAny", "count": 1 }, { "type": "ModelPreviewOverrideKJ", "count": 1 }, { "type": "bd04b316-09de-4e3a-b064-b13b79237cbd", "count": 1 }, { "type": "bc3d9d1b-4372-41dd-a0f1-83325748885d", "count": 1 }, { "type": "97ac977a-60fa-4a0f-a37c-36f79484a144", "count": 1 }, { "type": "OllamaOptionsV2", "count": 1 }, { "type": "OllamaGenerateV2", "count": 1 }, { "type": "OllamaConnectivityV2", "count": 1 }, { "type": "PrimitiveInt", "count": 1 }, { "type": "GetImageSize", "count": 1 }, { "type": "CustomCombo", "count": 1 }, { "type": "KSamplerSelect", "count": 1 }, { "type": "easy convertAnything", "count": 1 }, { "type": "RandomNoise", "count": 1 }, { "type": "BasicGuider", "count": 1 }, { "type": "SamplerCustomAdvanced", "count": 1 }, { "type": "VAEDecode", "count": 1 }, { "type": "VAEDecodeAudio", "count": 1 }, { "type": "MinimaxH3LatentUpscaler3D", "count": 1 }, { "type": "LTXVSeparateAVLatent", "count": 1 }, { "type": "LTXVConcatAVLatent", "count": 1 }, { "type": "MMH3TemporalSplitParamsV10", "count": 1 }, { "type": "MMH3SpatialSplitParamsV10", "count": 1 }, { "type": "ManualSigmas", "count": 1 }, { "type": "CLIPLoader", "count": 1 }, { "type": "MiniMaxH3MemoryEfficientSageAttentionPatch", "count": 1 }, { "type": "MiniMaxH3SigmaShift", "count": 1 }, { "type": "SpectrumApplyMiniMaxH3", "count": 1 }, { "type": "UNETLoader", "count": 1 }, { "type": "PathchSageAttentionKJ", "count": 1 }, { "type": "ModelPatchTorchSettings", "count": 1 }, { "type": "ModelAttentionBackend", "count": 1 }, { "type": "BlockSparseAttention", "count": 1 } ], "models": [ "H3\\minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors", "H3\\minimax_h3_ref_lora_rank_256_bf16.safetensors", "minimax_h3_audio_vae_fp32.safetensors", "minimax_h3_fl2va_pruned_int8_convrot.safetensors", "minimax_h3_latent_upscaler_3d_fp16.safetensors", "minimax_h3_video_vae_int8_convrot.safetensors", "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors", "taeh3.safetensors" ], "notes": [ "Video A, the generated cut, and Video B must all be the same size.", "Ref LoRA https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras # When using Ref LoRA, inference can be performed using the Fl2VA model. Controlnet Union https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/controlnet Comfyui_Minimax_h3_latent_Upscaler https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler Lightx2v Turbo-LoRA https://huggingface.co/lightx2v/Minimax-h3-Turbo Turbo-LoRA https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI PDD ACC LoRA https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras Spectrum-MiniMax-H3 https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3.git" ] }, { "file": "workflows/MiniMax_int8_FR-javano2609.1.1.json", "title": "Face Refine", "subtitle": "Automatic face detection and refinement pass", "description": "Refines faces in generated video with ComfyUI-H3-FaceRefine. For small faces, raise crop_size (e.g. 512-640) at the cost of slower processing.", "features": [ "FaceRefine", "Automatic detection", "INT8", "R2V pipeline", "Selectable crop size" ], "topLevelNodes": 18, "totalNodes": 78, "inventory": [ { "type": "ComfySwitchNode", "count": 6 }, { "type": "Note", "count": 4 }, { "type": "VHS_VideoCombine", "count": 3 }, { "type": "PrimitiveBoolean", "count": 3 }, { "type": "AddLabel", "count": 2 }, { "type": "VAELoader", "count": 2 }, { "type": "LoraLoaderModelOnly", "count": 2 }, { "type": "RandomNoise", "count": 2 }, { "type": "KSamplerSelect", "count": 2 }, { "type": "LTXVConcatAVLatent", "count": 2 }, { "type": "VAEEncodeAudio", "count": 2 }, { "type": "VAEDecode", "count": 2 }, { "type": "BasicScheduler", "count": 2 }, { "type": "MiniMaxH3ReferenceToVideo", "count": 2 }, { "type": "GetImageRangeFromBatch", "count": 2 }, { "type": "ComfyMathExpression", "count": 2 }, { "type": "2116b7a6-0103-4b9f-97d9-19d9554b768f", "count": 1 }, { "type": "ModelPreviewOverrideKJ", "count": 1 }, { "type": "b8839f65-32ab-4ce1-a3fc-4912368c2ed3", "count": 1 }, { "type": "0054af6e-df24-47e4-b99f-9e74d18a3519", "count": 1 }, { "type": "LoadVideoUI", "count": 1 }, { "type": "LoadImage", "count": 1 }, { "type": "2c9f2183-726b-494c-8073-4a254f5300f3", "count": 1 }, { "type": "87f3687c-de7e-48a5-af78-65788265c38d", "count": 1 }, { "type": "Group Controller", "count": 1 }, { "type": "ImageConcanate", "count": 1 }, { "type": "ImageScaleToMaxDimension", "count": 1 }, { "type": "ImageListToImageBatch", "count": 1 }, { "type": "GetImageSizeAndCount", "count": 1 }, { "type": "MiniMaxH3MemoryEfficientSageAttentionPatch", "count": 1 }, { "type": "CLIPLoader", "count": 1 }, { "type": "MiniMaxH3SigmaShift", "count": 1 }, { "type": "UNETLoader", "count": 1 }, { "type": "PathchSageAttentionKJ", "count": 1 }, { "type": "ModelPatchTorchSettings", "count": 1 }, { "type": "ModelAttentionBackend", "count": 1 }, { "type": "BlockSparseAttention", "count": 1 }, { "type": "SamplerCustomAdvanced", "count": 1 }, { "type": "H3FaceMaskSAM", "count": 1 }, { "type": "LTXVSeparateAVLatent", "count": 1 }, { "type": "H3InjectVideoLatent", "count": 1 }, { "type": "SAMLoader", "count": 1 }, { "type": "H3FaceStitch", "count": 1 }, { "type": "H3FaceTrackCrop", "count": 1 }, { "type": "H3PerFrameDenoise", "count": 1 }, { "type": "BasicGuider", "count": 1 }, { "type": "BatchImagesNode", "count": 1 }, { "type": "EmptyMiniMaxH3LatentAV", "count": 1 }, { "type": "GetLatentSizeAndCount", "count": 1 }, { "type": "VAEEncodeTiled", "count": 1 }, { "type": "MinimaxH3LatentUpscaler3D", "count": 1 }, { "type": "GetImageSize", "count": 1 }, { "type": "ManualSigmas", "count": 1 }, { "type": "MMH3SplitUpscale", "count": 1 } ], "models": [ "H3\\minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors", "H3\\minimax_h3_ref_lora_rank_256_bf16.safetensors", "minimax_h3_audio_vae_fp32.safetensors", "minimax_h3_fl2va_pruned_int8_convrot.safetensors", "minimax_h3_video_vae_int8_convrot.safetensors", "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors", "taeh3.safetensors" ], "notes": [ "Ref LoRA https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras FaceRefine https://github.com/Carasibana/ComfyUI-H3-FaceRefine Comfyui_Minimax_h3_latent_Upscaler https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler Face Model https://huggingface.co/Anzhc/Anzhcs_YOLOs https://huggingface.co/facebook/sam2.1-hiera-large", "If the face you want to refine is too small, increase \"crop_size\" to a value like 512 or 640; however, note that processing time will slow.", "Targets the range from 0 to the specified frame(range).", "Upscale Processing is very slow!" ] }, { "file": "workflows/MiniMax_int8_TTS-javano2609.1.json", "title": "Text-to-Speech / Video-to-Audio", "subtitle": "Fast audio-only speech generation, or new background audio for a mute clip", "description": "Generate speech from text or new ambient audio for an existing video. Not a lip-sync tool: it produces background sound, leaving the video frames untouched.", "features": [ "Audio-only", "TTS", "Video-to-Audio", "INT8", "SpeechLengthCalculator" ], "topLevelNodes": 9, "totalNodes": 47, "inventory": [ { "type": "ComfySwitchNode", "count": 7 }, { "type": "VAELoader", "count": 2 }, { "type": "LoraLoaderModelOnly", "count": 2 }, { "type": "ComfyMathExpression", "count": 2 }, { "type": "LoadImage", "count": 1 }, { "type": "SaveAudioAdvanced", "count": 1 }, { "type": "SpeechLengthCalculator", "count": 1 }, { "type": "466b2449-e267-4ec5-9751-96c224d14021", "count": 1 }, { "type": "VHS_LoadVideo", "count": 1 }, { "type": "VHS_VideoCombine", "count": 1 }, { "type": "1b57c27c-2176-4876-a458-298b3a1dd837", "count": 1 }, { "type": "PreviewAny", "count": 1 }, { "type": "Note", "count": 1 }, { "type": "RandomNoise", "count": 1 }, { "type": "BasicGuider", "count": 1 }, { "type": "KSamplerSelect", "count": 1 }, { "type": "SamplerCustomAdvanced", "count": 1 }, { "type": "CLIPLoader", "count": 1 }, { "type": "VAEDecodeAudio", "count": 1 }, { "type": "BasicScheduler", "count": 1 }, { "type": "easy int", "count": 1 }, { "type": "MiniMaxH3SigmaShift", "count": 1 }, { "type": "UNETLoader", "count": 1 }, { "type": "ModelPatchTorchSettings", "count": 1 }, { "type": "MiniMaxH3MemoryEfficientSageAttentionPatch", "count": 1 }, { "type": "PathchSageAttentionKJ", "count": 1 }, { "type": "ModelAttentionBackend", "count": 1 }, { "type": "BlockSparseAttention", "count": 1 }, { "type": "LTXVSeparateAVLatent", "count": 1 }, { "type": "VAEEncode", "count": 1 }, { "type": "LTXVConcatAVLatent", "count": 1 }, { "type": "ResizeImageMaskNode", "count": 1 }, { "type": "MiniMaxH3ReferenceToVideo", "count": 1 }, { "type": "GetImageSize", "count": 1 }, { "type": "PrimitiveInt", "count": 1 }, { "type": "OllamaOptionsV2", "count": 1 }, { "type": "OllamaGenerateV2", "count": 1 }, { "type": "OllamaConnectivityV2", "count": 1 } ], "models": [ "H3\\minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors", "H3\\minimax_h3_ref_lora_rank_256_bf16.safetensors", "minimax_h3_audio_vae_fp32.safetensors", "minimax_h3_fl2va_pruned_int8_convrot.safetensors", "minimax_h3_video_vae_int8_convrot.safetensors", "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors" ], "notes": [ "Video-to-Audio does not regenerate the video. In other words, it is not intended for lip sync. Its purpose is to generate new background sounds that are missing from the video." ] }, { "file": "extra/MiniMax-Music3-javano2608.1.json", "title": "Music 3", "subtitle": "Songs up to 5 minutes from a structured music caption", "description": "MiniMax Music 3 generates complete, stable songs (up to 5 min) from a music description with an optional reference audio track.", "features": [ "Music generation", "Up to 5 min", "Caption + reference audio" ], "topLevelNodes": 7, "totalNodes": 22, "inventory": [ { "type": "PreviewAny", "count": 1 }, { "type": "PreviewAudio", "count": 1 }, { "type": "MarkdownNote", "count": 1 }, { "type": "ac99f841-a3de-4329-9564-953b81cf9e16", "count": 1 }, { "type": "SaveAudioAdvanced", "count": 1 }, { "type": "7ef02436-c462-4950-94cb-2d3358359c65", "count": 1 }, { "type": "PrimitiveStringMultiline", "count": 1 }, { "type": "CLIPLoader", "count": 1 }, { "type": "VAELoader", "count": 1 }, { "type": "ConditioningZeroOut", "count": 1 }, { "type": "EmptyMiniMaxMusic3LatentAudio", "count": 1 }, { "type": "KSampler", "count": 1 }, { "type": "SeedNode", "count": 1 }, { "type": "VAEDecodeAudio", "count": 1 }, { "type": "ComfySwitchNode", "count": 1 }, { "type": "AudioEnhancementNode", "count": 1 }, { "type": "Reroute", "count": 1 }, { "type": "MiniMaxMusic3TextEncode", "count": 1 }, { "type": "DiffusionModelLoaderKJ", "count": 1 }, { "type": "OllamaOptionsV2", "count": 1 }, { "type": "OllamaConnectivityV2", "count": 1 }, { "type": "OllamaGenerateV2", "count": 1 } ], "models": [ "minimax_music3_dav.safetensors", "minimax_music3_dit_fp32.safetensors", "minimax_music3_text_encoder_pruned_bf16.safetensors" ], "notes": [ "## MiniMax Music 3 A music generation model by MiniMax that creates complete songs up to 5 minutes long with stable structure and high audio quality. Two inputs drive it: **Caption** (structured music description: style, mood, vocals, arrangement) and **Lyrics** (with `[intro]` `[verse]` `[chorus]` `[bridge]` `[outro]` tags controlling song structure). ### Parameters | Parameter | Description | |---|---| | Caption | Music description. Write it in three sections — Global Metadata → Vocal Details → Arrangement. The more specific, the closer the result. | | Lyrics | Song lyrics + section tags. Tags are the only executable structural instructions; the lyric text itself only conveys mood. | | max_duration | Target song length in seconds (default 120 = 2 minutes; the model supports up to ~300 s / 5 minutes). Longer songs take more time and VRAM. | | seed | Random seed for the generation. Keep it fixed to reproduce the same song; change it to get a different take. | | Model files | diffusion model / text encoder / VAE.| | Tiled decode | Decode the audio VAE in overlapping tiles to drastically cut VRAM usage — helpful for long songs on low-VRAM GPUs. Slightly slower, with a small risk of seams at tile boundaries; turn it off on high-VRAM GPUs for the best quality. | ### Prompt Writing Guide You can find an official [Music Caption Rewriter Skill here](https://github.com/MiniMax-AI/MiniMax-Music3)" ] }, { "file": "extra/MiniMax_int8-Ref2Image-javano2608.2.json", "title": "Ref2Image", "subtitle": "Reference-guided single image generation", "description": "INT8 image generation conditioned on a reference image and Ref LoRA, with the Latent Upscaler for high-resolution output.", "features": [ "Reference image", "INT8", "Ref LoRA", "Latent Upscaler" ], "topLevelNodes": 15, "totalNodes": 51, "inventory": [ { "type": "LoraLoaderModelOnly", "count": 7 }, { "type": "ComfySwitchNode", "count": 5 }, { "type": "LoadImage", "count": 4 }, { "type": "Note", "count": 2 }, { "type": "VAELoader", "count": 2 }, { "type": "CFGGuider", "count": 2 }, { "type": "SamplerCustomAdvanced", "count": 2 }, { "type": "BasicScheduler", "count": 2 }, { "type": "9122ef1f-6c5c-4bbb-bfb1-88492720488e", "count": 1 }, { "type": "PreviewAny", "count": 1 }, { "type": "MarkdownNote", "count": 1 }, { "type": "ResolutionSelector", "count": 1 }, { "type": "ModelPreviewOverrideKJ", "count": 1 }, { "type": "d82a9cc9-ee6f-47c5-9e22-6fcdc0e7146a", "count": 1 }, { "type": "4b173715-6297-498a-a94b-2e5aed2a91ea", "count": 1 }, { "type": "SaveImageAdvanced", "count": 1 }, { "type": "85362315-ca4f-4a6f-b770-afd3b5a06baa", "count": 1 }, { "type": "MiniMaxH3MemoryEfficientSageAttentionPatch", "count": 1 }, { "type": "DiffusionModelLoaderKJ", "count": 1 }, { "type": "MiniMaxH3SigmaShift", "count": 1 }, { "type": "CLIPLoader", "count": 1 }, { "type": "OllamaOptionsV2", "count": 1 }, { "type": "OllamaGenerateV2", "count": 1 }, { "type": "OllamaConnectivityV2", "count": 1 }, { "type": "KSamplerSelect", "count": 1 }, { "type": "ConditioningZeroOut", "count": 1 }, { "type": "VAEDecode", "count": 1 }, { "type": "ImageFromBatch", "count": 1 }, { "type": "RandomNoise", "count": 1 }, { "type": "MiniMaxH3ReferenceToVideo", "count": 1 }, { "type": "LTXVSeparateAVLatent", "count": 1 }, { "type": "MinimaxH3LatentUpscaler3D", "count": 1 }, { "type": "LTXVConcatAVLatent", "count": 1 } ], "models": [ "H3\\minimax_h3_ref2v_lightx2v_turbo_4step_v0.1_resized_avg_rank_20_bf16.safetensors", "H3\\minimax_h3_ref_lora_rank_256_bf16.safetensors", "minimax_h3_audio_vae_fp32.safetensors", "minimax_h3_fl2va_pruned_int8_convrot.safetensors", "minimax_h3_latent_upscaler_3d_fp16.safetensors", "minimax_h3_video_vae_int8_convrot.safetensors", "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors", "taeh3.safetensors" ], "notes": [ "| megapixels | Aspect | Output (multiple=32) | |---|---|---| | 0.2 | 16:9 | 608 x 352 | | 0.3 | 16:9 | 736 x 416 | | 0.4 | 16:9 | 864 x 480 | | 0.5 | 16:9 | 960 x 544 | | 0.6 | 16:9 | 1056 x 608 | | 0.7 | 16:9 | 1152 x 640 | | 0.8 | 16:9 | 1216 x 672 | | 0.9 | 16:9 | 1280 x 736 | | 0.98 | 16:9 | 1344 x 768 | | 1.0 | 16:9 | 1376 x 768 | | 1.2 | 16:9 | 1504 x 832 | | 1.5 | 16:9 | 1664 x 928 | | 1.8 | 16:9 | 1824 x 1024 | | 2.0 | 16:9 | 1920 x 1088 |", "Ref LoRA https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras Comfyui_Minimax_h3_latent_Upscaler https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler" ] } ] };