H3_Easy_Ref2V_Workflow / Nugget_H3_EasyR2V_Prompter_Only_v03.json
PoopMan333's picture
Upload 45 files
2d75cf2 verified
Raw
History Blame Contribute Delete
45 kB
{"id":"2eec6263-88ce-462b-80d4-90475d887f69","revision":0,"last_node_id":144,"last_link_id":199,"nodes":[{"id":86,"type":"LoadImage","pos":[-1090,2830],"size":[230,330],"flags":{},"order":0,"mode":0,"inputs":[{"localized_name":"image","name":"image","type":"COMBO","widget":{"name":"image"},"link":null},{"localized_name":"choose file to upload","name":"upload","type":"IMAGEUPLOAD","widget":{"name":"upload"},"link":null}],"outputs":[{"localized_name":"IMAGE","name":"IMAGE","type":"IMAGE","links":[193]},{"localized_name":"MASK","name":"MASK","type":"MASK","links":null}],"title":"Load image 1","properties":{"Node name for S&R":"LoadImage","cnr_id":"comfy-core","ver":"0.21.1"},"widgets_values":["ChatGPT Image Aug 29, 2026, 05_55_00 AM.png","image"],"widgets_values_named":{"image":"ChatGPT Image Aug 29, 2026, 05_55_00 AM.png","upload":"image"},"color":"#222","bgcolor":"#000"},{"id":99,"type":"LoadImage","pos":[-820,2840],"size":[230,330],"flags":{},"order":1,"mode":0,"inputs":[{"localized_name":"image","name":"image","type":"COMBO","widget":{"name":"image"},"link":null},{"localized_name":"choose file to upload","name":"upload","type":"IMAGEUPLOAD","widget":{"name":"upload"},"link":null}],"outputs":[{"localized_name":"IMAGE","name":"IMAGE","type":"IMAGE","links":[194]},{"localized_name":"MASK","name":"MASK","type":"MASK","links":null}],"title":"Load image 2","properties":{"Node name for S&R":"LoadImage","cnr_id":"comfy-core","ver":"0.21.1"},"widgets_values":["RidersKirby.webp","image"],"widgets_values_named":{"image":"RidersKirby.webp","upload":"image"},"color":"#222","bgcolor":"#000"},{"id":89,"type":"PreviewAny","pos":[520,2220],"size":[540,570],"flags":{"collapsed":false},"order":14,"mode":0,"inputs":[{"localized_name":"source","name":"source","type":"*","link":196}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":null}],"title":"Enhanced prompt","properties":{"Node name for S&R":"PreviewAny","cnr_id":"comfy-core","ver":"0.16.4"},"widgets_values":[],"widgets_values_named":{},"color":"#323","bgcolor":"#535"},{"id":91,"type":"PrimitiveStringMultiline","pos":[-1660,2350],"size":[450,320],"flags":{},"order":3,"mode":0,"inputs":[{"localized_name":"value","name":"value","type":"STRING","widget":{"name":"value"},"link":null}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":[101]}],"title":"User prompt","properties":{"Node name for S&R":"PrimitiveStringMultiline","cnr_id":"comfy-core","ver":"0.30.0"},"widgets_values":["the dolphin from picture 1 is fighting kirby from picture 2. write a very creative scene. "],"widgets_values_named":{"value":"the dolphin from picture 1 is fighting kirby from picture 2. write a very creative scene. "},"color":"#232","bgcolor":"#353"},{"id":136,"type":"PrimitiveStringMultiline","pos":[-1090,1740],"size":[310,360],"flags":{},"order":4,"mode":0,"inputs":[{"localized_name":"value","name":"value","type":"STRING","widget":{"name":"value"},"link":null}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":null}],"title":"Notes","properties":{"Node name for S&R":"PrimitiveStringMultiline"},"widgets_values":["VIDEO AUTO TRANSCRIBER\n\nprompt max length - 0 follows description detail. Raise only if\n descriptions stop mid-sentence; the console warns when they do.\ndescription detail - how much is written per shot, and how many frames\n the model sees. Frame count scales with shot length.\nscene sensitivity - how hard it looks for cuts. Raise if shots are\n missed, lower if one shot is split in two.\nmin shot - shortest allowed shot. Two cuts closer than this become one,\n so a high value throws cuts away. Fast-cut music video: 0.5.\n Interviews: 5.\nmax frame size - a ceiling, not a target. A 640-wide clip stays 640 even\n at 1024. Lower it for speed on less VRAM.\nwhisper model - large-v3 is most accurate. Set to 'off' for clips with\n no dialogue: it skips audio entirely and saves real time.\naudio language - named languages, not codes. Leave on auto.\n\nOUTPUTS\n\nimages / audio / fps / frame_count replace a Get Video Components node.\nfull description is the whole block. overview, characters identified\nand shots are the same content split up, so a downstream node can take\njust one part without parsing it back out.\naudio transcription is the dialogue on its own with timings.\naudio language is the detected name, for tagging <d>[English] ...</d>."],"widgets_values_named":{"value":"VIDEO AUTO TRANSCRIBER\n\nprompt max length - 0 follows description detail. Raise only if\n descriptions stop mid-sentence; the console warns when they do.\ndescription detail - how much is written per shot, and how many frames\n the model sees. Frame count scales with shot length.\nscene sensitivity - how hard it looks for cuts. Raise if shots are\n missed, lower if one shot is split in two.\nmin shot - shortest allowed shot. Two cuts closer than this become one,\n so a high value throws cuts away. Fast-cut music video: 0.5.\n Interviews: 5.\nmax frame size - a ceiling, not a target. A 640-wide clip stays 640 even\n at 1024. Lower it for speed on less VRAM.\nwhisper model - large-v3 is most accurate. Set to 'off' for clips with\n no dialogue: it skips audio entirely and saves real time.\naudio language - named languages, not codes. Leave on auto.\n\nOUTPUTS\n\nimages / audio / fps / frame_count replace a Get Video Components node.\nfull description is the whole block. overview, characters identified\nand shots are the same content split up, so a downstream node can take\njust one part without parsing it back out.\naudio transcription is the dialogue on its own with timings.\naudio language is the detected name, for tagging <d>[English] ...</d>."},"color":"#432","bgcolor":"#653"},{"id":138,"type":"PrimitiveStringMultiline","pos":[90,1740],"size":[290,360],"flags":{},"order":5,"mode":0,"inputs":[{"localized_name":"value","name":"value","type":"STRING","widget":{"name":"value"},"link":null}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":null}],"title":"Notes","properties":{"Node name for S&R":"PrimitiveStringMultiline"},"widgets_values":["GENERATE PROMPT (WITH REFERENCES)\n\nReplaces the built-in Generate Text plus an image batching node.\nImages keep their own aspect ratios and are never cropped.\n\nimg description detail - appends an instruction to describe the\n reference images and raises the token budget. Leave OFF when the\n system prompt already specifies its own format, or the two fight.\n Use it when you want this node as a plain captioner.\nmax image size - longest side each reference is scaled to. A cap, not a\n target: smaller images are left alone, never upscaled.\nmax length - token budget for the reply. A full six-section H3 prompt is\n around 650 tokens, so 1024 leaves headroom. A budget you never reach\n costs nothing.\nsampling - off is greedy: faster, and the same input gives the same\n output so you can tell whether a change helped. Turning it on reveals\n temperature and repetition penalty.\n\nModel notes: works with Gemma 4 12B uncensored and Qwen3-VL 8B. The 8B is\nfaster at transcription and slower at prompt writing, but easier on VRAM.\nOthers untested."],"widgets_values_named":{"value":"GENERATE PROMPT (WITH REFERENCES)\n\nReplaces the built-in Generate Text plus an image batching node.\nImages keep their own aspect ratios and are never cropped.\n\nimg description detail - appends an instruction to describe the\n reference images and raises the token budget. Leave OFF when the\n system prompt already specifies its own format, or the two fight.\n Use it when you want this node as a plain captioner.\nmax image size - longest side each reference is scaled to. A cap, not a\n target: smaller images are left alone, never upscaled.\nmax length - token budget for the reply. A full six-section H3 prompt is\n around 650 tokens, so 1024 leaves headroom. A budget you never reach\n costs nothing.\nsampling - off is greedy: faster, and the same input gives the same\n output so you can tell whether a change helped. Turning it on reveals\n temperature and repetition penalty.\n\nModel notes: works with Gemma 4 12B uncensored and Qwen3-VL 8B. The 8B is\nfaster at transcription and slower at prompt writing, but easier on VRAM.\nOthers untested."},"color":"#432","bgcolor":"#653"},{"id":93,"type":"JoinStringMulti","pos":[-250,2460],"size":[310,190],"flags":{},"order":12,"mode":0,"inputs":[{"localized_name":"string_1","name":"string_1","type":"STRING","link":101},{"localized_name":"string_2","name":"string_2","shape":7,"type":"STRING","link":190},{"localized_name":"inputcount","name":"inputcount","type":"INT","widget":{"name":"inputcount"},"link":null},{"localized_name":"delimiter","name":"delimiter","type":"STRING","widget":{"name":"delimiter"},"link":null},{"localized_name":"return_list","name":"return_list","type":"BOOLEAN","widget":{"name":"return_list"},"link":null},{"name":"string_3","shape":7,"type":"STRING","link":191}],"outputs":[{"localized_name":"string","name":"string","type":"STRING","links":[195]}],"title":"Join string (user + system prompt)","properties":{"Node name for S&R":"JoinStringMulti","cnr_id":"comfyui-kjnodes","ver":"35e5956193769d18a13136cdedb73a36a05c73e6"},"widgets_values":[3,"",false,null],"widgets_values_named":{"inputcount":3,"delimiter":"","return_list":false,"Update inputs":null},"color":"#222","bgcolor":"#000"},{"id":141,"type":"PrimitiveStringMultiline","pos":[-210,2720],"size":[560,200],"flags":{"collapsed":true},"order":6,"mode":0,"inputs":[{"localized_name":"value","name":"value","type":"STRING","widget":{"name":"value"},"link":null}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":[191]}],"title":"System prompt - v7.1 mini","properties":{"Node name for S&R":"PrimitiveStringMultiline","cnr_id":"comfy-core","ver":"0.30.0"},"widgets_values":["MINIMAX H3 REF2V PROMPT GUIDE v7\n\n# ROLE\nYou write prompts only for MiniMax H3 Reference-to-Video (Ref2V).\nConvert up to 9 ordered reference images, up to 1 source video with transcript, and the user's instruction into one precise Ref2V prompt.\nThe user instruction controls the requested result. References control the visible, temporal, or audible properties to preserve, transfer, replace, or reuse.\nPlan silently. Output only the finished prompt.\n\n# LANGUAGE\nWrite all six sections in English. Preserve the original language only for dialogue or lyrics inside <d> and for visible on-screen text.\n\n# REFERENCES\nImages retain their supplied order: <Picture 1> through <Picture 9>. Never renumber or invent references.\n\nUse <Subject N> for independently reusable visible content, including people, animals, objects, clothing, props, environments, styles, poses, actions, expressions, interfaces, and effects.\n- One image may define multiple subjects.\n- Multiple references may define one subject; state what each contributes.\n- Do not invent details that are not visible.\n- Keep each label's meaning stable.\n\nUse standalone <Picture N> only when the image itself is a concrete first frame, keyframe, last frame, edited frame, storyboard, or composition anchor. Otherwise cite the picture inside the relevant <Subject N> definition.\n\nWhen a source video is directly edited or supplies whole-video structure, define:\n<Video 1> is the source video for the target edit.\nUse <Video 1> for camera, cuts, timing, staging, continuation, or overall temporal structure. Visible content reused from it should be defined separately as <Subject N> when it must be tracked.\n\nWhen enabled source audio is copied or independently referenced, define <Audio 1>. Do not create <Audio 1> merely because the video contains sound.\n\n# EDITING AND RETENTION\nExplicitly state what is retained and what changes. The user's requested change overrides the references.\n\nFor replacement, identify the original subject and replacement clearly:\nReplace the original performer with <Subject 1>, transferring <Subject 1>'s referenced appearance while preserving the relevant motion, timing, gestures, expression, framing, and prop interaction from <Video 1>.\nAdapt motion naturally. Track multiple replacements independently. State shot or timestamp boundaries for partial or mid-video replacements.\n\nVisible-reference markers:\n- fully_preserved: the defined properties remain\n- partially_preserved: only specified properties remain or some change\n- attribute_transfer: referenced properties transfer to another identifiable target\n- weak_reference: only broad style, category, composition, or mood guides the result\n\nAudio-reference markers:\n- fully_copy: complete source audio becomes the complete final track\n- partially_copy: only part or selected layers are copied, or copied audio is modified\n- reference: timbre, delivery, rhythm, dialogue content, music style, or sound texture guides new audio without copying the signal\n- weak_reference: only broad audio character or atmosphere is retained\n\nDo not promise perfect identity, pixel-perfect reproduction, or exact physical accuracy.\n\n# MOTION, SHOTS, AND CONTINUITY\nDescribe the target in playback order using observable instructions: composition, subject appearance and position, environment, lighting, action and state changes, camera behavior, current sound, and when each reference takes effect.\nPreserve source camera work, performance timing, cuts, environment, and continuity unless the user requests changes.\nDo not add unnecessary cuts, camera motion, characters, props, dialogue, or plot events.\nObjects must not teleport, duplicate, disappear, or change without cause.\n\nUse:\n[Shot 1] ...\n[Shot 2] At MM:SS.mmm, ...\nThe opening shot normally has no timestamp. Use timestamps for later cuts or precise changes.\nIf the user specifies a target duration, do not describe scenes, dialogue, audio, or events beyond that time, even if the source video or transcript continues.\n\n# DIALOGUE AND AUDIO\nOriginal audio reuse: define <Audio 1>, state that it is copied from the source, and use the correct audio retention marker.\nIf music alone changes, retain requested dialogue and non-music audio while describing the new score in non_diegetic_music.\n\nReplacement dialogue:\n- preserve the requested words exactly\n- identify the speaker and reuse stable speaker IDs in order of first vocal event\n- synchronize mouth movement\n- retain source timing, tone, expression, or delivery when requested\n\nFormat:\n<Subject N> (S1) says: <d>[Language] Dialogue.</d>\nUse [unclear] for unintelligible speech. Do not guess.\nDescribe dialogue, singing, synchronized sounds, and active <Audio N> relationships in the relevant shot. Do not repeat full dialogue in the sound sections.\nDo not invent dialogue unless requested.\n\n# REQUIRED OUTPUT\nOutput exactly these six sections in this order:\nsubject_definitions:\nsummary:\nretention_analysis:\ndetailed_description:\noverall_soundscape:\nnon_diegetic_music:\n\nsubject_definitions:\nDefine only references actually used. Give each tracked item one line with its source, role, and key properties.\n\nsummary:\nWrite one short paragraph beginning with the applicable combined task labels:\n[keyframe completion], [reference generation], [video editing], [video continuation], [audio reuse], [audio reference]\nJoin applicable labels with \" + \". Do not add a type merely because that media exists.\nFor a direct source edit, begin after the label with: The target video is an edited version of <Video 1>.\n\nretention_analysis:\nGive one line per defined label:\n<Subject N> (appears in [Shot ...]): marker - exact retained or transferred properties.\n<Picture N> ([Shot ...] frame role): marker - exact frame or composition relationship.\n<Video 1> (relevant structural properties): marker - exact retained or referenced properties.\n<Audio 1>: audio_marker - exact copied or referenced audio properties.\nUse only properties relevant to the result. Do not include speaker IDs here.\n\ndetailed_description:\nThis is the main generation instruction. Start with at most two sentences establishing overall appearance and source continuity when useful, then describe every shot in playback order. Explicitly map replacements, actions, expressions, framing, camera, timing, environment, lighting, props, sound, dialogue, lip sync, continuity, and each shot's final state where relevant. Avoid plot-only summaries and generic quality terms.\n\noverall_soundscape:\nSummarize ambience and physical sounds across the video. State applicable <Audio N> copy/reference behavior. Keep shot-synchronized events in detailed_description.\n\nnon_diegetic_music:\nDescribe audience-only music, including style/instrumentation, tempo, and development, only when requested or present through audio reuse/reference. Otherwise write N/A.\n\n# SILENT CHECK\nBefore output, verify:\n1. Exactly six sections, in order, with no introduction, notes, markdown fence, or closing text.\n2. Every label comes from supplied media, remains stable, is used, and has a retention line.\n3. Replacements and their timing are explicit.\n4. User-requested changes override source properties; other relevant source performance and structure remain.\n5. Audio copying versus audio reference is distinguished correctly.\n6. Dialogue is exact, attributed, language-tagged, and lip-synchronized when relevant.\n7. detailed_description contains actual visual, temporal, and audible instructions.\n8. Nothing unrelated is invented.\n"],"widgets_values_named":{"value":"MINIMAX H3 REF2V PROMPT GUIDE v7\n\n# ROLE\nYou write prompts only for MiniMax H3 Reference-to-Video (Ref2V).\nConvert up to 9 ordered reference images, up to 1 source video with transcript, and the user's instruction into one precise Ref2V prompt.\nThe user instruction controls the requested result. References control the visible, temporal, or audible properties to preserve, transfer, replace, or reuse.\nPlan silently. Output only the finished prompt.\n\n# LANGUAGE\nWrite all six sections in English. Preserve the original language only for dialogue or lyrics inside <d> and for visible on-screen text.\n\n# REFERENCES\nImages retain their supplied order: <Picture 1> through <Picture 9>. Never renumber or invent references.\n\nUse <Subject N> for independently reusable visible content, including people, animals, objects, clothing, props, environments, styles, poses, actions, expressions, interfaces, and effects.\n- One image may define multiple subjects.\n- Multiple references may define one subject; state what each contributes.\n- Do not invent details that are not visible.\n- Keep each label's meaning stable.\n\nUse standalone <Picture N> only when the image itself is a concrete first frame, keyframe, last frame, edited frame, storyboard, or composition anchor. Otherwise cite the picture inside the relevant <Subject N> definition.\n\nWhen a source video is directly edited or supplies whole-video structure, define:\n<Video 1> is the source video for the target edit.\nUse <Video 1> for camera, cuts, timing, staging, continuation, or overall temporal structure. Visible content reused from it should be defined separately as <Subject N> when it must be tracked.\n\nWhen enabled source audio is copied or independently referenced, define <Audio 1>. Do not create <Audio 1> merely because the video contains sound.\n\n# EDITING AND RETENTION\nExplicitly state what is retained and what changes. The user's requested change overrides the references.\n\nFor replacement, identify the original subject and replacement clearly:\nReplace the original performer with <Subject 1>, transferring <Subject 1>'s referenced appearance while preserving the relevant motion, timing, gestures, expression, framing, and prop interaction from <Video 1>.\nAdapt motion naturally. Track multiple replacements independently. State shot or timestamp boundaries for partial or mid-video replacements.\n\nVisible-reference markers:\n- fully_preserved: the defined properties remain\n- partially_preserved: only specified properties remain or some change\n- attribute_transfer: referenced properties transfer to another identifiable target\n- weak_reference: only broad style, category, composition, or mood guides the result\n\nAudio-reference markers:\n- fully_copy: complete source audio becomes the complete final track\n- partially_copy: only part or selected layers are copied, or copied audio is modified\n- reference: timbre, delivery, rhythm, dialogue content, music style, or sound texture guides new audio without copying the signal\n- weak_reference: only broad audio character or atmosphere is retained\n\nDo not promise perfect identity, pixel-perfect reproduction, or exact physical accuracy.\n\n# MOTION, SHOTS, AND CONTINUITY\nDescribe the target in playback order using observable instructions: composition, subject appearance and position, environment, lighting, action and state changes, camera behavior, current sound, and when each reference takes effect.\nPreserve source camera work, performance timing, cuts, environment, and continuity unless the user requests changes.\nDo not add unnecessary cuts, camera motion, characters, props, dialogue, or plot events.\nObjects must not teleport, duplicate, disappear, or change without cause.\n\nUse:\n[Shot 1] ...\n[Shot 2] At MM:SS.mmm, ...\nThe opening shot normally has no timestamp. Use timestamps for later cuts or precise changes.\nIf the user specifies a target duration, do not describe scenes, dialogue, audio, or events beyond that time, even if the source video or transcript continues.\n\n# DIALOGUE AND AUDIO\nOriginal audio reuse: define <Audio 1>, state that it is copied from the source, and use the correct audio retention marker.\nIf music alone changes, retain requested dialogue and non-music audio while describing the new score in non_diegetic_music.\n\nReplacement dialogue:\n- preserve the requested words exactly\n- identify the speaker and reuse stable speaker IDs in order of first vocal event\n- synchronize mouth movement\n- retain source timing, tone, expression, or delivery when requested\n\nFormat:\n<Subject N> (S1) says: <d>[Language] Dialogue.</d>\nUse [unclear] for unintelligible speech. Do not guess.\nDescribe dialogue, singing, synchronized sounds, and active <Audio N> relationships in the relevant shot. Do not repeat full dialogue in the sound sections.\nDo not invent dialogue unless requested.\n\n# REQUIRED OUTPUT\nOutput exactly these six sections in this order:\nsubject_definitions:\nsummary:\nretention_analysis:\ndetailed_description:\noverall_soundscape:\nnon_diegetic_music:\n\nsubject_definitions:\nDefine only references actually used. Give each tracked item one line with its source, role, and key properties.\n\nsummary:\nWrite one short paragraph beginning with the applicable combined task labels:\n[keyframe completion], [reference generation], [video editing], [video continuation], [audio reuse], [audio reference]\nJoin applicable labels with \" + \". Do not add a type merely because that media exists.\nFor a direct source edit, begin after the label with: The target video is an edited version of <Video 1>.\n\nretention_analysis:\nGive one line per defined label:\n<Subject N> (appears in [Shot ...]): marker - exact retained or transferred properties.\n<Picture N> ([Shot ...] frame role): marker - exact frame or composition relationship.\n<Video 1> (relevant structural properties): marker - exact retained or referenced properties.\n<Audio 1>: audio_marker - exact copied or referenced audio properties.\nUse only properties relevant to the result. Do not include speaker IDs here.\n\ndetailed_description:\nThis is the main generation instruction. Start with at most two sentences establishing overall appearance and source continuity when useful, then describe every shot in playback order. Explicitly map replacements, actions, expressions, framing, camera, timing, environment, lighting, props, sound, dialogue, lip sync, continuity, and each shot's final state where relevant. Avoid plot-only summaries and generic quality terms.\n\noverall_soundscape:\nSummarize ambience and physical sounds across the video. State applicable <Audio N> copy/reference behavior. Keep shot-synchronized events in detailed_description.\n\nnon_diegetic_music:\nDescribe audience-only music, including style/instrumentation, tempo, and development, only when requested or present through audio reuse/reference. Otherwise write N/A.\n\n# SILENT CHECK\nBefore output, verify:\n1. Exactly six sections, in order, with no introduction, notes, markdown fence, or closing text.\n2. Every label comes from supplied media, remains stable, is used, and has a retention line.\n3. Replacements and their timing are explicit.\n4. User-requested changes override source properties; other relevant source performance and structure remain.\n5. Audio copying versus audio reference is distinguished correctly.\n6. Dialogue is exact, attributed, language-tagged, and lip-synchronized when relevant.\n7. detailed_description contains actual visual, temporal, and audible instructions.\n8. Nothing unrelated is invented.\n"},"color":"#232","bgcolor":"#353"},{"id":142,"type":"NuggetGeneratePrompt","pos":[100,2230],"size":[268.9859375,292],"flags":{},"order":13,"mode":0,"inputs":[{"localized_name":"clip","name":"clip","type":"CLIP","link":192},{"localized_name":"image_1","name":"image_1","shape":7,"type":"IMAGE","link":193},{"localized_name":"image_2","name":"image_2","shape":7,"type":"IMAGE","link":194},{"localized_name":"image_3","name":"image_3","shape":7,"type":"IMAGE","link":null},{"localized_name":"img_description_detail","name":"img_description_detail","type":"COMBO","widget":{"name":"img_description_detail"},"link":null},{"localized_name":"prompt","name":"prompt","type":"STRING","widget":{"name":"prompt"},"link":195},{"localized_name":"max_image_size","name":"max_image_size","type":"COMBO","widget":{"name":"max_image_size"},"link":null},{"localized_name":"max_length","name":"max_length","type":"INT","widget":{"name":"max_length"},"link":null},{"localized_name":"sampling_mode","name":"sampling_mode","type":"COMFY_DYNAMICCOMBO_V3","widget":{"name":"sampling_mode"},"link":null},{"localized_name":"seed","name":"seed","shape":7,"type":"INT","widget":{"name":"seed"},"link":null}],"outputs":[{"localized_name":"generated_text","name":"generated_text","type":"STRING","links":[196]}],"properties":{"Node name for S&R":"NuggetGeneratePrompt","cnr_id":"comfyui-nugget","ver":"1.1.0"},"widgets_values":["normal","",1024,1024,"off",0,"randomize"],"widgets_values_named":{"img_description_detail":"normal","prompt":"","max_image_size":1024,"max_length":1024,"sampling_mode":"off","seed":0,"control_after_generate":"randomize"}},{"id":85,"type":"CLIPLoader","pos":[-1650,2160],"size":[430,130],"flags":{},"order":7,"mode":0,"inputs":[{"localized_name":"clip_name","name":"clip_name","type":"COMBO","widget":{"name":"clip_name"},"link":null},{"localized_name":"type","name":"type","type":"COMBO","widget":{"name":"type"},"link":null},{"localized_name":"device","name":"device","shape":7,"type":"COMBO","widget":{"name":"device"},"link":null}],"outputs":[{"localized_name":"CLIP","name":"CLIP","type":"CLIP","links":[192,197]}],"properties":{"Node name for S&R":"CLIPLoader","cnr_id":"comfy-core","ver":"0.16.4"},"widgets_values":["qwen3vl_8b_fp8_scaled.safetensors","minimax","default"],"widgets_values_named":{"clip_name":"qwen3vl_8b_fp8_scaled.safetensors","type":"minimax","device":"default"},"color":"#222","bgcolor":"#000"},{"id":130,"type":"PreviewAny","pos":[-710,2220],"size":[290,300],"flags":{"collapsed":false},"order":11,"mode":0,"inputs":[{"localized_name":"source","name":"source","type":"*","link":199}],"outputs":[{"localized_name":"STRING","name":"STRING","type":"STRING","links":[190]}],"title":"Source video description","properties":{"Node name for S&R":"PreviewAny","cnr_id":"comfy-core","ver":"0.16.4"},"widgets_values":[],"widgets_values_named":{},"color":"#222","bgcolor":"#000"},{"id":143,"type":"Note","pos":[-700,2570],"size":[280,110],"flags":{},"order":8,"mode":0,"inputs":[],"outputs":[],"properties":{},"widgets_values":["If you want to transcribe audio, please go to \"custom_nodes/comfyui-nugget/\" and go to the \"install.bat\" or \"install.py\". This will install Whisper which is required for audio, otherwise can leave OFF\n"],"widgets_values_named":{"text":"If you want to transcribe audio, please go to \"custom_nodes/comfyui-nugget/\" and go to the \"install.bat\" or \"install.py\". This will install Whisper which is required for audio, otherwise can leave OFF\n"},"color":"#432","bgcolor":"#653"},{"id":104,"type":"LoadVideo","pos":[-1600,2810],"size":[440,360],"flags":{"collapsed":false},"order":9,"mode":0,"inputs":[{"localized_name":"file","name":"file","type":"COMBO","widget":{"name":"file"},"link":null},{"localized_name":"choose file to upload","name":"upload","type":"IMAGEUPLOAD","widget":{"name":"upload"},"link":null}],"outputs":[{"localized_name":"VIDEO","name":"VIDEO","type":"VIDEO","links":[198]}],"title":"Load source video","properties":{"Node name for S&R":"LoadVideo","cnr_id":"comfy-core","ver":"0.30.0"},"widgets_values":["Singin' in the Rain 14s.mp4","image"],"widgets_values_named":{"file":"Singin' in the Rain 14s.mp4","upload":"image"},"color":"#432","bgcolor":"#653"},{"id":144,"type":"VideoAutoTranscribe","pos":[-1070,2210],"size":[310,510],"flags":{},"order":10,"mode":0,"inputs":[{"localized_name":"clip","name":"clip","type":"CLIP","link":197},{"localized_name":"video","name":"video","shape":7,"type":"VIDEO","link":198},{"localized_name":"mode","name":"mode","type":"COMBO","widget":{"name":"mode"},"link":null},{"localized_name":"prompt_max_length","name":"prompt_max_length","type":"INT","widget":{"name":"prompt_max_length"},"link":null},{"localized_name":"description_detail","name":"description_detail","type":"COMBO","widget":{"name":"description_detail"},"link":null},{"localized_name":"scene_sensitivity","name":"scene_sensitivity","type":"COMBO","widget":{"name":"scene_sensitivity"},"link":null},{"localized_name":"min_shot","name":"min_shot","type":"FLOAT","widget":{"name":"min_shot"},"link":null},{"localized_name":"whisper_model","name":"whisper_model","type":"COMBO","widget":{"name":"whisper_model"},"link":null},{"localized_name":"audio_language","name":"audio_language","type":"COMBO","widget":{"name":"audio_language"},"link":null},{"localized_name":"frame_layout","name":"frame_layout","shape":7,"type":"COMBO","widget":{"name":"frame_layout"},"link":null},{"localized_name":"max_frame_size","name":"max_frame_size","shape":7,"type":"COMBO","widget":{"name":"max_frame_size"},"link":null},{"localized_name":"seed","name":"seed","shape":7,"type":"INT","widget":{"name":"seed"},"link":null},{"localized_name":"show","name":"show","shape":7,"type":"COMBO","widget":{"name":"show"},"link":null}],"outputs":[{"localized_name":"video_images","name":"video_images","type":"IMAGE","links":null},{"localized_name":"video audio","name":"video audio","type":"AUDIO","links":null},{"localized_name":"fps","name":"fps","type":"FLOAT","links":null},{"localized_name":"frame_count","name":"frame_count","type":"INT","links":null},{"localized_name":"full description","name":"full description","type":"STRING","links":[199]}],"properties":{"Node name for S&R":"VideoAutoTranscribe"},"widgets_values":["General",1600,"normal","normal",1,"large-v3","auto","grid",768,45156651615615620,"fixed","compact"],"widgets_values_named":{"mode":"General","prompt_max_length":1600,"description_detail":"normal","scene_sensitivity":"normal","min_shot":1,"whisper_model":"large-v3","audio_language":"auto","frame_layout":"grid","max_frame_size":768,"seed":45156651615615620,"control_after_generate":"fixed","show":"compact"}},{"id":97,"type":"MarkdownNote","pos":[-2300,2190],"size":[420,680],"flags":{},"order":2,"mode":0,"inputs":[],"outputs":[],"title":"START HERE - Downloads & Setup","properties":{},"widgets_values":["# Nugget H3 EasyR2V β€” Prompter Only\n\n**By [C_Nugget](https://huggingface.co/PoopMan333)** β€” models, LoRAs and other bits over on the HuggingFace.\n\n\n**[If this has helped you, consider sending through any tips of a few dollars, any tips helps with the power bills. Thank you.](https://ko-fi.com/c_nugget)**\n\n---\n\n## How it works\n\nThis graph **writes H3 prompts**. It does not generate video β€” feed the output into an H3 workflow, or copy it out by hand.\n\n1. **Video Auto Transcriber** watches a reference clip (optional) and writes a plain-English description β€” shots, camera moves, dialogue, sound.\n2. **Generate prompt (with references)** takes that description, your own instruction and up to 9 reference images, and hands them to a vision-language model that writes the finished H3 prompt.\n3. The enhanced prompt appears in the preview on the right. Copy it, or wire it straight into an H3 graph.\n\nWork down this list once. After that, the graph runs left to right.\n\n---\n\n## 1. Install the custom node packs\n\nVia **ComfyUI Manager**, or `git clone` into `ComfyUI/custom_nodes/`:\n\n| pack | what this graph uses it for |\n|---|---|\n| **ComfyUI-Nugget** | `Video Auto Transcriber`, `Generate prompt (with references)` |\n| **ComfyUI-KJNodes** | `JoinStringMulti` |\n\n**If the file will not open, a pack is missing.** ComfyUI is refusing a node type it does not recognise. Bypassing will not help β€” install the pack, restart, then reopen.\n\n---\n\n## 2. Run the Nugget setup script\n\n**Not optional if you want dialogue transcribed.**\n\n```\nComfyUI/custom_nodes/ComfyUI-Nugget/install.bat\n```\n\nDouble-click it, or from a terminal:\n\n```\npython install.py\n```\n\nIt installs `faster-whisper`, then **builds a real model and transcribes a second of silence** to prove it works β€” importing the package proves nothing, because the CUDA libraries only load when a model is constructed.\n\n**Do not just run `pip install faster-whisper`.** ComfyUI runs its own Python; a normal terminal installs into a different one, and the node will still report it missing. The script finds the right interpreter and re-runs itself there.\n\n`python install.py --check` reports without changing anything.\n\nWithout it the transcriber still describes the video β€” you simply get no transcript.\n\n---\n\n## 3. The language model\n\nOne model does everything here: it reads the video and writes the prompt. Goes in `models/text_encoders/`, then pick it in **Load CLIP**.\n\n| model | size | notes |\n|---|---|---|\n| [Qwen3-VL 8B fp8_scaled](https://huggingface.co/Comfy-Org/Qwen3-VL/tree/main/text_encoders) | 10.6 GB | **Start here.** Currently selected. Best all-round vision + prompt writing. |\n| [Qwen3-VL 8B nvfp4](https://huggingface.co/Comfy-Org/Qwen3-VL/tree/main/text_encoders) | 6.3 GB | Same model, smaller quant. Use on 12 GB cards. |\n| [Gemma 4 E4B fp8](https://huggingface.co/Comfy-Org/gemma-4/tree/main/text_encoders) | 9.1 GB | ComfyUI's original default. Fine if you already have it. |\n| [DeepNeuralNerd Gemma4 12B heretic](https://huggingface.co/DeepNeuralNerd/Gemma-4-12B-it-uncensored-heretic-DeepNeuralNerd-LTX_2.5_ComfyUI) | ~12 GB | Uncensored. Wants 20 GB+ in practice. |\n\nAnything else is untested.\n\n---\n\n## 4. Run order\n\n1. **Image input** β€” 1 to 9 reference pictures. Sockets appear as you fill them.\n2. **Use Video Transcription** β€” optional. Load a clip and it is described automatically, cuts and camera movement included.\n3. **H3 prompt enhancer** β€” joins your instruction, the system prompt and the video description, then writes the finished H3 prompt.\n\nThe enhanced prompt appears in the preview on the right. Copy it, or wire it into an H3 graph.\n\n**To skip the video**, bypass **Load source video** (Ctrl+B). The transcriber returns empty strings and costs nothing β€” the prompt is written from the images and your text alone.\n\n---\n\n## 5. Re-running cheaply\n\nComfyUI skips any node whose inputs have not changed. Once a clip is transcribed, **queue again** and the transcriber is cached and free β€” only what you actually changed re-runs.\n\nKeep both `seed` widgets on **fixed**. Set one to `randomize` and its inputs change every queue, which is a cache miss β€” you re-transcribe the whole video every run for nothing.\n\nOn the prompt node, `randomize` only produces *different* text when **sampling** is on. With sampling off it re-runs for an identical answer.\n\n---\n\n## 6. If something goes wrong\n\n| symptom | cause |\n|---|---|\n| The file will not open | A node pack from section 1 is missing. |\n| `'SD1ClipModel' object has no attribute 'generate'` | Load CLIP is loading a Stable Diffusion encoder, not a language model. |\n| ComfyUI **exits** with a code, no error | Out of VRAM. Check the `GB free before loading the text encoder` line in the console and use a smaller model (drop from fp8_scaled to nvfp4). |\n| No transcript, says faster-whisper missing | It went into the wrong Python. Run `install.bat` from section 2. |\n| Prompt stops mid-sentence | Raise **max length** on the prompt node. The console warns when it happens. |\n| Dialogue invented over music | Set **whisper model** to `off`. `large-v3` hallucinates on non-speech audio. |\n| A reference image seems to be ignored | Bypassing a Load Image leaves the link but sends nothing, and the rest renumber β€” `image_4` becomes picture 2. Disconnect instead of bypassing. The console says when it happens. |\n\nClips run best around 8–20 seconds. The transcriber returns every frame, so a minute of 1080p is more RAM than most machines want to spend. Trim first.\n\n---\n\n## 7. Credit\n\nWorkflow build, tuning and prompt-enhancer chain by **C_Nugget** β€” [huggingface.co/PoopMan333](https://huggingface.co/PoopMan333).\n\nOriginal single-image workflow by **mackyb** (H3 Basic prompt enhancer v2). This is a modified variant, not their release.\n"],"widgets_values_named":{"text":"# Nugget H3 EasyR2V β€” Prompter Only\n\n**By [C_Nugget](https://huggingface.co/PoopMan333)** β€” models, LoRAs and other bits over on the HuggingFace.\n\n\n**[If this has helped you, consider sending through any tips of a few dollars, any tips helps with the power bills. Thank you.](https://ko-fi.com/c_nugget)**\n\n---\n\n## How it works\n\nThis graph **writes H3 prompts**. It does not generate video β€” feed the output into an H3 workflow, or copy it out by hand.\n\n1. **Video Auto Transcriber** watches a reference clip (optional) and writes a plain-English description β€” shots, camera moves, dialogue, sound.\n2. **Generate prompt (with references)** takes that description, your own instruction and up to 9 reference images, and hands them to a vision-language model that writes the finished H3 prompt.\n3. The enhanced prompt appears in the preview on the right. Copy it, or wire it straight into an H3 graph.\n\nWork down this list once. After that, the graph runs left to right.\n\n---\n\n## 1. Install the custom node packs\n\nVia **ComfyUI Manager**, or `git clone` into `ComfyUI/custom_nodes/`:\n\n| pack | what this graph uses it for |\n|---|---|\n| **ComfyUI-Nugget** | `Video Auto Transcriber`, `Generate prompt (with references)` |\n| **ComfyUI-KJNodes** | `JoinStringMulti` |\n\n**If the file will not open, a pack is missing.** ComfyUI is refusing a node type it does not recognise. Bypassing will not help β€” install the pack, restart, then reopen.\n\n---\n\n## 2. Run the Nugget setup script\n\n**Not optional if you want dialogue transcribed.**\n\n```\nComfyUI/custom_nodes/ComfyUI-Nugget/install.bat\n```\n\nDouble-click it, or from a terminal:\n\n```\npython install.py\n```\n\nIt installs `faster-whisper`, then **builds a real model and transcribes a second of silence** to prove it works β€” importing the package proves nothing, because the CUDA libraries only load when a model is constructed.\n\n**Do not just run `pip install faster-whisper`.** ComfyUI runs its own Python; a normal terminal installs into a different one, and the node will still report it missing. The script finds the right interpreter and re-runs itself there.\n\n`python install.py --check` reports without changing anything.\n\nWithout it the transcriber still describes the video β€” you simply get no transcript.\n\n---\n\n## 3. The language model\n\nOne model does everything here: it reads the video and writes the prompt. Goes in `models/text_encoders/`, then pick it in **Load CLIP**.\n\n| model | size | notes |\n|---|---|---|\n| [Qwen3-VL 8B fp8_scaled](https://huggingface.co/Comfy-Org/Qwen3-VL/tree/main/text_encoders) | 10.6 GB | **Start here.** Currently selected. Best all-round vision + prompt writing. |\n| [Qwen3-VL 8B nvfp4](https://huggingface.co/Comfy-Org/Qwen3-VL/tree/main/text_encoders) | 6.3 GB | Same model, smaller quant. Use on 12 GB cards. |\n| [Gemma 4 E4B fp8](https://huggingface.co/Comfy-Org/gemma-4/tree/main/text_encoders) | 9.1 GB | ComfyUI's original default. Fine if you already have it. |\n| [DeepNeuralNerd Gemma4 12B heretic](https://huggingface.co/DeepNeuralNerd/Gemma-4-12B-it-uncensored-heretic-DeepNeuralNerd-LTX_2.5_ComfyUI) | ~12 GB | Uncensored. Wants 20 GB+ in practice. |\n\nAnything else is untested.\n\n---\n\n## 4. Run order\n\n1. **Image input** β€” 1 to 9 reference pictures. Sockets appear as you fill them.\n2. **Use Video Transcription** β€” optional. Load a clip and it is described automatically, cuts and camera movement included.\n3. **H3 prompt enhancer** β€” joins your instruction, the system prompt and the video description, then writes the finished H3 prompt.\n\nThe enhanced prompt appears in the preview on the right. Copy it, or wire it into an H3 graph.\n\n**To skip the video**, bypass **Load source video** (Ctrl+B). The transcriber returns empty strings and costs nothing β€” the prompt is written from the images and your text alone.\n\n---\n\n## 5. Re-running cheaply\n\nComfyUI skips any node whose inputs have not changed. Once a clip is transcribed, **queue again** and the transcriber is cached and free β€” only what you actually changed re-runs.\n\nKeep both `seed` widgets on **fixed**. Set one to `randomize` and its inputs change every queue, which is a cache miss β€” you re-transcribe the whole video every run for nothing.\n\nOn the prompt node, `randomize` only produces *different* text when **sampling** is on. With sampling off it re-runs for an identical answer.\n\n---\n\n## 6. If something goes wrong\n\n| symptom | cause |\n|---|---|\n| The file will not open | A node pack from section 1 is missing. |\n| `'SD1ClipModel' object has no attribute 'generate'` | Load CLIP is loading a Stable Diffusion encoder, not a language model. |\n| ComfyUI **exits** with a code, no error | Out of VRAM. Check the `GB free before loading the text encoder` line in the console and use a smaller model (drop from fp8_scaled to nvfp4). |\n| No transcript, says faster-whisper missing | It went into the wrong Python. Run `install.bat` from section 2. |\n| Prompt stops mid-sentence | Raise **max length** on the prompt node. The console warns when it happens. |\n| Dialogue invented over music | Set **whisper model** to `off`. `large-v3` hallucinates on non-speech audio. |\n| A reference image seems to be ignored | Bypassing a Load Image leaves the link but sends nothing, and the rest renumber β€” `image_4` becomes picture 2. Disconnect instead of bypassing. The console says when it happens. |\n\nClips run best around 8–20 seconds. The transcriber returns every frame, so a minute of 1080p is more RAM than most machines want to spend. Trim first.\n\n---\n\n## 7. Credit\n\nWorkflow build, tuning and prompt-enhancer chain by **C_Nugget** β€” [huggingface.co/PoopMan333](https://huggingface.co/PoopMan333).\n\nOriginal single-image workflow by **mackyb** (H3 Basic prompt enhancer v2). This is a modified variant, not their release.\n"},"color":"#222","bgcolor":"#000"}],"links":[[101,91,0,93,0,"STRING"],[190,130,0,93,1,"STRING"],[191,141,0,93,5,"STRING"],[192,85,0,142,0,"CLIP"],[193,86,0,142,1,"IMAGE"],[194,99,0,142,2,"IMAGE"],[195,93,0,142,5,"STRING"],[196,142,0,89,0,"STRING"],[197,85,0,144,0,"CLIP"],[198,104,0,144,1,"VIDEO"],[199,144,4,130,0,"STRING"]],"groups":[{"id":1,"title":"H3 prompt enhancer - multi-image (1-9)","bounding":[-300,2130,750,790],"color":"#3f789e","flags":{}},{"id":2,"title":"Image input","bounding":[-1650,2710,1315.8841858676296,796.5907307396988],"color":"#3f789e","flags":{}},{"id":3,"title":"Use Video Transcription","bounding":[-1160,2130,821.1533827183462,547.692634049236],"color":"#3f789e","flags":{}}],"config":{},"extra":{"ds":{"scale":1.2754897606277609,"offset":[2608.125955876652,-2098.291530021562]},"frontendVersion":"1.49.6","VHS_latentpreview":false,"VHS_latentpreviewrate":0,"VHS_MetadataImage":true,"VHS_KeepIntermediate":true},"version":0.4}