{"id":"2eec6263-88ce-462b-80d4-90475d887f69","revision":0,"last_node_id":144,"last_link_id":199,"nodes":[{"id":86,"type":"LoadImage","pos":[-1090,2830],"size":[230,330],"flags":{},"order":0,"mode":0,"inputs":[{"localized_name":"image","name":"image","type":"COMBO","widget":{"name":"image"},"link":null},{"localized_name":"choose file to upload","name":"upload","type":"IMAGEUPLOAD","widget":{"name":"upload"},"link":null}],"outputs":[{"localized_name":"IMAGE","name":"IMAGE","type":"IMAGE","links":[193]},{"localized_name":"MASK","name":"MASK","type":"MASK","links":null}],"title":"Load image 1","properties":{"Node name for S&R":"LoadImage","cnr_id":"comfy-core","ver":"0.21.1"},"widgets_values":["ChatGPT Image Aug 29, 2026, 05_55_00 AM.png","image"],"widgets_values_named":{"image":"ChatGPT Image Aug 29, 2026, 05_55_00 AM.png","upload":"image"},"color":"#222","bgcolor":"#000"},{"id":99,"type":"LoadImage","pos":[-820,2840],"size":[230,330],"flags":{},"order":1,"mode":0,"inputs":[{"localized_name":"image","name":"image","type":"COMBO","widget":{"name":"image"},"link":null},{"localized_name":"choose file to upload","name":"upload","type":"IMAGEUPLOAD","widget":{"name":"upload"},"link":null}],"outputs":[{"localized_name":"IMAGE","name":"IMAGE","type":"IMAGE","links":[194]},{"localized_name":"MASK","name":"MASK","type":"MASK","links":null}],"title":"Load image 2","properties":{"Node name for S&R":"LoadImage","cnr_id":"comfy-core","ver":"0.21.1"},"widgets_values":["RidersKirby.webp","image"],"widgets_values_named":{"image":"RidersKirby.webp","upload":"image"},"color":"#222","bgcolor":"#000"},{"id":89,"type":"PreviewAny","pos":[520,2220],"size":[540,570],"flags":{"collapsed":false},"order":14,"mode":0,"inputs":[{"localized_name":"source","name":"source","type":"*","link":196}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":null}],"title":"Enhanced prompt","properties":{"Node name for S&R":"PreviewAny","cnr_id":"comfy-core","ver":"0.16.4"},"widgets_values":[],"widgets_values_named":{},"color":"#323","bgcolor":"#535"},{"id":91,"type":"PrimitiveStringMultiline","pos":[-1660,2350],"size":[450,320],"flags":{},"order":3,"mode":0,"inputs":[{"localized_name":"value","name":"value","type":"STRING","widget":{"name":"value"},"link":null}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":[101]}],"title":"User prompt","properties":{"Node name for S&R":"PrimitiveStringMultiline","cnr_id":"comfy-core","ver":"0.30.0"},"widgets_values":["the dolphin from picture 1 is fighting kirby from picture 2. write a very creative scene. "],"widgets_values_named":{"value":"the dolphin from picture 1 is fighting kirby from picture 2. write a very creative scene. "},"color":"#232","bgcolor":"#353"},{"id":136,"type":"PrimitiveStringMultiline","pos":[-1090,1740],"size":[310,360],"flags":{},"order":4,"mode":0,"inputs":[{"localized_name":"value","name":"value","type":"STRING","widget":{"name":"value"},"link":null}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":null}],"title":"Notes","properties":{"Node name for S&R":"PrimitiveStringMultiline"},"widgets_values":["VIDEO AUTO TRANSCRIBER\n\nprompt max length - 0 follows description detail. Raise only if\n descriptions stop mid-sentence; the console warns when they do.\ndescription detail - how much is written per shot, and how many frames\n the model sees. Frame count scales with shot length.\nscene sensitivity - how hard it looks for cuts. Raise if shots are\n missed, lower if one shot is split in two.\nmin shot - shortest allowed shot. Two cuts closer than this become one,\n so a high value throws cuts away. Fast-cut music video: 0.5.\n Interviews: 5.\nmax frame size - a ceiling, not a target. A 640-wide clip stays 640 even\n at 1024. Lower it for speed on less VRAM.\nwhisper model - large-v3 is most accurate. Set to 'off' for clips with\n no dialogue: it skips audio entirely and saves real time.\naudio language - named languages, not codes. Leave on auto.\n\nOUTPUTS\n\nimages / audio / fps / frame_count replace a Get Video Components node.\nfull description is the whole block. overview, characters identified\nand shots are the same content split up, so a downstream node can take\njust one part without parsing it back out.\naudio transcription is the dialogue on its own with timings.\naudio language is the detected name, for tagging [English] ...."],"widgets_values_named":{"value":"VIDEO AUTO TRANSCRIBER\n\nprompt max length - 0 follows description detail. Raise only if\n descriptions stop mid-sentence; the console warns when they do.\ndescription detail - how much is written per shot, and how many frames\n the model sees. Frame count scales with shot length.\nscene sensitivity - how hard it looks for cuts. Raise if shots are\n missed, lower if one shot is split in two.\nmin shot - shortest allowed shot. Two cuts closer than this become one,\n so a high value throws cuts away. Fast-cut music video: 0.5.\n Interviews: 5.\nmax frame size - a ceiling, not a target. A 640-wide clip stays 640 even\n at 1024. Lower it for speed on less VRAM.\nwhisper model - large-v3 is most accurate. Set to 'off' for clips with\n no dialogue: it skips audio entirely and saves real time.\naudio language - named languages, not codes. Leave on auto.\n\nOUTPUTS\n\nimages / audio / fps / frame_count replace a Get Video Components node.\nfull description is the whole block. overview, characters identified\nand shots are the same content split up, so a downstream node can take\njust one part without parsing it back out.\naudio transcription is the dialogue on its own with timings.\naudio language is the detected name, for tagging [English] ...."},"color":"#432","bgcolor":"#653"},{"id":138,"type":"PrimitiveStringMultiline","pos":[90,1740],"size":[290,360],"flags":{},"order":5,"mode":0,"inputs":[{"localized_name":"value","name":"value","type":"STRING","widget":{"name":"value"},"link":null}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":null}],"title":"Notes","properties":{"Node name for S&R":"PrimitiveStringMultiline"},"widgets_values":["GENERATE PROMPT (WITH REFERENCES)\n\nReplaces the built-in Generate Text plus an image batching node.\nImages keep their own aspect ratios and are never cropped.\n\nimg description detail - appends an instruction to describe the\n reference images and raises the token budget. Leave OFF when the\n system prompt already specifies its own format, or the two fight.\n Use it when you want this node as a plain captioner.\nmax image size - longest side each reference is scaled to. A cap, not a\n target: smaller images are left alone, never upscaled.\nmax length - token budget for the reply. A full six-section H3 prompt is\n around 650 tokens, so 1024 leaves headroom. A budget you never reach\n costs nothing.\nsampling - off is greedy: faster, and the same input gives the same\n output so you can tell whether a change helped. Turning it on reveals\n temperature and repetition penalty.\n\nModel notes: works with Gemma 4 12B uncensored and Qwen3-VL 8B. The 8B is\nfaster at transcription and slower at prompt writing, but easier on VRAM.\nOthers untested."],"widgets_values_named":{"value":"GENERATE PROMPT (WITH REFERENCES)\n\nReplaces the built-in Generate Text plus an image batching node.\nImages keep their own aspect ratios and are never cropped.\n\nimg description detail - appends an instruction to describe the\n reference images and raises the token budget. Leave OFF when the\n system prompt already specifies its own format, or the two fight.\n Use it when you want this node as a plain captioner.\nmax image size - longest side each reference is scaled to. A cap, not a\n target: smaller images are left alone, never upscaled.\nmax length - token budget for the reply. A full six-section H3 prompt is\n around 650 tokens, so 1024 leaves headroom. A budget you never reach\n costs nothing.\nsampling - off is greedy: faster, and the same input gives the same\n output so you can tell whether a change helped. Turning it on reveals\n temperature and repetition penalty.\n\nModel notes: works with Gemma 4 12B uncensored and Qwen3-VL 8B. The 8B is\nfaster at transcription and slower at prompt writing, but easier on VRAM.\nOthers untested."},"color":"#432","bgcolor":"#653"},{"id":93,"type":"JoinStringMulti","pos":[-250,2460],"size":[310,190],"flags":{},"order":12,"mode":0,"inputs":[{"localized_name":"string_1","name":"string_1","type":"STRING","link":101},{"localized_name":"string_2","name":"string_2","shape":7,"type":"STRING","link":190},{"localized_name":"inputcount","name":"inputcount","type":"INT","widget":{"name":"inputcount"},"link":null},{"localized_name":"delimiter","name":"delimiter","type":"STRING","widget":{"name":"delimiter"},"link":null},{"localized_name":"return_list","name":"return_list","type":"BOOLEAN","widget":{"name":"return_list"},"link":null},{"name":"string_3","shape":7,"type":"STRING","link":191}],"outputs":[{"localized_name":"string","name":"string","type":"STRING","links":[195]}],"title":"Join string (user + system prompt)","properties":{"Node name for S&R":"JoinStringMulti","cnr_id":"comfyui-kjnodes","ver":"35e5956193769d18a13136cdedb73a36a05c73e6"},"widgets_values":[3,"",false,null],"widgets_values_named":{"inputcount":3,"delimiter":"","return_list":false,"Update inputs":null},"color":"#222","bgcolor":"#000"},{"id":141,"type":"PrimitiveStringMultiline","pos":[-210,2720],"size":[560,200],"flags":{"collapsed":true},"order":6,"mode":0,"inputs":[{"localized_name":"value","name":"value","type":"STRING","widget":{"name":"value"},"link":null}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":[191]}],"title":"System prompt - v7.1 mini","properties":{"Node name for S&R":"PrimitiveStringMultiline","cnr_id":"comfy-core","ver":"0.30.0"},"widgets_values":["MINIMAX H3 REF2V PROMPT GUIDE v7\n\n# ROLE\nYou write prompts only for MiniMax H3 Reference-to-Video (Ref2V).\nConvert up to 9 ordered reference images, up to 1 source video with transcript, and the user's instruction into one precise Ref2V prompt.\nThe user instruction controls the requested result. References control the visible, temporal, or audible properties to preserve, transfer, replace, or reuse.\nPlan silently. Output only the finished prompt.\n\n# LANGUAGE\nWrite all six sections in English. Preserve the original language only for dialogue or lyrics inside and for visible on-screen text.\n\n# REFERENCES\nImages retain their supplied order: through . Never renumber or invent references.\n\nUse for independently reusable visible content, including people, animals, objects, clothing, props, environments, styles, poses, actions, expressions, interfaces, and effects.\n- One image may define multiple subjects.\n- Multiple references may define one subject; state what each contributes.\n- Do not invent details that are not visible.\n- Keep each label's meaning stable.\n\nUse standalone only when the image itself is a concrete first frame, keyframe, last frame, edited frame, storyboard, or composition anchor. Otherwise cite the picture inside the relevant definition.\n\nWhen a source video is directly edited or supplies whole-video structure, define:\n