Workflow: V2V - Extend Any Video
** V2V - Extend Any Video**
Take an input video, and extend it seamlessly to a longer one.
The workflow lets you set where you want to continue from, and how long you want the extension to be
This first workflow uses latent masking to seamlessly extend.
Will also add a variant that simply uses prompt (works, but more sensitive to correct prompt), as well as a H3 Guider node variant.
Feel free to try it out here:
https://huggingface.co/RuneXX/Minimax-H3-Workflows/tree/main/Video-to-Video
I've tried every extension workflow so far with minimax and yours is the cleanest and easiest to use. Always prefer your workflow's over the rest. I am curious though, were does one get the nodes for the chunk feed forward and low vram attention? I already have kijai SolAttn triton and Saganaki22 sol-attn but it has a different chunkfeedforward node then yours and I can't find any github page for the lowvram node which I would very much like to try. Any chance of a link of where I can get it?
chunk feed forwards and low vram should both be part of KJ Nodes (from Kijai). Make sure its up to date.
I should have mentioned that in the post ; -)
Thank you Rune. I keep forgetting to update those nodes along with Comfy. Kijai is always adding great new stuff.
Also forgot to mention, this workflow is awesome. Thanks so much Rune. Like I stated earlier, I tried every extended workflow out there and this bury's them all. its working incredibly well. Thank you again. I check your hugging face page everyday for any updates from you. So glad I didn't stop after the slowdown with LTX.
Happy to hear ;-)
And thats the goal, just to make something user-friendly and share, as well as try different ways to use the model ; - )
No slowdown (hopefully hehe).. . just a bit of summer days being a bit less on the computer ;- )
Glad to hear it. I need to get out more as well. I was afraid you were maybe losing interest.
Its also great to see your messing around with Minimax now. Its so much better then LTX for keeping the likeness of your characters. I just wish it wasn't so darn slow. Hopefully that Fal is legit and releases the super fast model and there is no quality loss. Anyway, thanks again Rune for all your workflows. You really know your stuff when it comes to making these.
By the way, you may want to mention on your LTX 2.3 page that you now have a MiniMax workflow page. I'm sure theres a lot of people on there that may not be aware.
Hello! Thanks for the awesome workflow! What is the difference between Duration and Extend Duration? The latter is the length of the extension itself, but what is Duration for?
Hello! Thanks for the awesome workflow! What is the difference between Duration and Extend Duration? The latter is the length of the extension itself, but what is Duration for?
Good question ;- ) Will check, but probably an oversight, a left over from workflow generating video.
Its the Extend Duration that is in use ;-)
Will update the workflow soon, so I'll remove the redundant one. The updated workflow has a frame blend for difficult videos where Minimax cant quite match the color space (as well as an optional color match node).
For most videos Minimax will do fine though ;-)
Thank you! I have been learning a lot about comfyui/workflows in general with your workflows.
It seems that using the generated prompt is a must to be able to extend it seamlessly otherwise it doesn't quite continue from the last frame of the original video.
I may suggest to add node to view the generated prompt. In my tests, sometimes it will add actions like "leaning forward" that I did not specify in the original prompt.
So you can edit the generated prompt if needed, copy it, bypass enhanced prompt and run it again.
But so far, I still cant get it to be at the same quality of my extended videos generated with LTX. I am guessing is because of the turbo Lora.
Yes this one is all in latent, so it the prompt matters.
I will update the workflow with a enhanced prompt preview (also added a few blend frames for difficult source videos)
Will also make 2 other variants of extend video that Minimax can do:
- extend with guider (this one can be a bit easier to prompt for)
- extend only with prompt (its a bit "difficult one", as in the prompt as to be very specific. But i'll add a note for a skeleton prompt)
Very nicely done 😃
Reminds me very much of my vace workflows with extending.
I noticed that the browser (Chrome on windows 11) somehow becomes non-responsive lately and I have to refresh the page regularly which can mess up the queued items sometimes. This happens when I use your otherwise amazing V2V - Extend Any Video workflow. Did anyone else experience this or I customized it too much for my needs and it confuses the browser? :-) When I switch to other workflows, this does not happen.
I noticed that the browser (Chrome on windows 11) somehow becomes non-responsive lately and I have to refresh the page regularly which can mess up the queued items sometimes. This happens when I use your otherwise amazing V2V - Extend Any Video workflow. Did anyone else experience this or I customized it too much for my needs and it confuses the browser? :-) When I switch to other workflows, this does not happen.
This IMO is a comfy thing... it happens MUCH more than it used to ever since the comfy kitchen attention patch they did. I think they have been doing things with the memory use that causes problems. I use firefox and have the same issues with any flows now. Even if my vram is not maxed. I run a 5060 16gb and it sits at 11 and firefox becomes a stuttering mess while on that tab.
I did find it worked better with the
Modern Node Design (Nodes 2.0) enabled. I hate Nodes 2.0.... but i think its the way they are leaning now so the old way may be becoming more and more glitchy. It does not fix it but you can scroll around easier with it enabled.
It's an awesome workflow. I don't really have the words to describe it. It is my absolute favourite. :-)
Would it be difficult to use this latent masking with a latent loaded from a file which was previously saved? That way the vae decode/encode process could be skipped and the image quality degradation may be a lot less after several chaining steps.
I think that should work. Let me know if it solves the contrast shifting? I tried everything. I have my own version of this now but it shifts, i dont know if its the model or the video compression over time. Upscaling batched works quite well tbh
Does it happen without turbo/distill lora or using a different one of them?
I'm not that good in comfyUI to update his workflow with latent loading and processing. :-(
Im good with comfy, i tried like 15 different ways to do latent stuff... nothing worked to eliminate the shift. There may be a way, i just havent found it yet.
I can say upscaling batched works. If you can make a 30 second video with minimax at 0.2 or 0.4 you can upscale it batched without much issues if you have the pc to do it.
These are my flows, im working on this daily to make this like i used to use vace.
https://civitai.red/models/2847150/minimax-simple
There is no way to avoid degradation without camera cuts. Anyone with some mystery node or workflow telling you this is possible, is lying to you.
Fortunately, this workflow works great with cuts. You can do a camera change mid sentence/action and it will continue the next part of the speech/action smoothly without any notice of skips in the video or audio. After a lot of testing of getting this workflow to work right, I actually prefer the camera cuts now instead of just one long boring single camera shot. I make 20 to 30 second clips and do a camera change at the start of every new clip. Works wonderfully.
Word of advice though, remove the prompt enhancer. it messes things up royally and is useless. Also, set up all your character images and backgrounds as subjects in your prompt and refer to them only as subjects for speech and actions. this includes added voices. Doing this fixes any audio issues or character drift. Also stick with the FL2VA model. Even though its not meant for references, it works better for me for some odd reason. It follows the prompts better and has better image quality and keeps the face likeness from images a lot better. Its also faster and uses less vram. I've made full 5 minute videos without any degradation or skips in video/audio with this workflow.
This workflow does mess up sometimes on the previous videos audio. Even though the audio was fine before loading the video in to this workflow, sometimes when the join happens it messes up the voices or music causing cracking or missing words near the end of the old video. I can't figure out how to avoid this or fix it. Fortunately Rune included a second saved file of just the extended part that you can join up to the old file which only requires one or two frames to be deleted with a video editor to make it seamless.
Hey Rune, any clue why this happens about 50% of the time with a join? Changing the overlap doesn't fix it although I notice that with a higher number then 2, it does seem to happen a lot more.
Also, with this workflow, is it possible to take the MiniMax H3 Reference to Video node and place it in the main screen without having to carry over the rest of the nodes inside the subgraph of V2V - Extend? I like the clean format of your workflow but would really like the Reference to Video node showing on the main screen. Just a minor inconvenience of having to go into the subgraph and change pics/audio files all the time.
The audio problem i believe is due to video helper suite video loader duration output being a float but only to one decimal place while the rest of the nodes all use 3 decimal places. This causes an audio skip if the audio is not exactly to one decimal place. To get around this i used mtb nodes audio duration node. It outputs the full millisecond length of the audio and you can convert that to 3 decimal floats.
Thanks for confirming that nothing avoids the degradation.
I thought this may work but i never tested it. I dont know if using latents to continue over RGB is better. My flows use rgb because cutting latents in the wrong spot causes the surrounding latents to go all messed up.
https://github.com/chanon/comfyui-obvpm-timeline
This seems useful but not sure if it solves the contrast shift or degradation
I found the best way to get a 30 second continuus shot is to do the shot at 0.2-0.6mp fully. Minimax can do 30-40 seconds and remain proper. More causes a loop of some kind. Then you can upscale this easily batched at 1080p in batches of 124 frames or even less. I upscale 0.2mp videos to 2mp in one stage of 2 steps... Latent upscaling is crap and slow.
Yep, I just started using that workflow shortly after making my previous post. Its really well done and you can make 3 or 4 extensions without having to do a cut with that one. I love Runes workflows and have always stuck with them but I have to be honest here, that one you linked to is the best one I've ever used for extending. I highly recommend it to everyone now. I appreciate your advice on how to fix the audio skips but I guess I don't really need to mess with that now on account of obvpm.
By the way, I don't use any turbo loras when making long videos. Even with cuts, those babies can really mess things up bad. Stick to Spectrum node and Comfy Kitchen Or Sage for speedups.
Nice so you think it solves the extensions drift? I thought it might but have not had the time to test it. Its like a perfect little node, exactly how i would have done a full workflow just to accomplish.
I will need to try it 😃
I use the 3 step lora from
https://huggingface.co/TaoLiveAIGC/TaoMate-H3/tree/main
I know its text to video lora but i use it for ref to video
I use it at 0.8 and use 4 steps. It does really well for that and it does 2 step upscales. I make 0.2-0.4mp videos at 10-35 seconds then upscale via batching with 2 steps... it can actually upscale fully from 0.2mp to 2.0mp. I think my upscale still does a better job than any latent upscale.
I can do full 1080p 30 second videos with it in 4 steps + 2 for upscale. Its quite fast.
I just wish we had a true controlnet like vace for minimax. Its so close but fun control is messy.
Unfortunately it still drifts which I knew it would, despite the comments on that channel. This is something I don't think will ever be resolved. You can extend it longer before it starts to deteriorate compared to Runes. You can get about 3 to 4 extensions done before you should really do a cut to prevent this from happening. At least with that workflow, it doesn't require as many cuts.
As for your upscaler, I would love to try it but I don't think my puny 5070 12gb can handle it. I don't bother using the obvpm Upscaler. As you stated above, those take just way too damn long. If I need a high resolution vid, Ill just generate at 1280x736 15 seconds and extend from there. I really don't see much difference between 1280 up to 1920. Anything above 1280 I believe gives diminishing returns. Or maybe its just my old tired eyes.
As for that lora you suggested, I already tried it. I've tried them all. The one you recommended has the best audio output over the rest but it's terrible at following my prompts and misses so many little details in my prompts compared to not using anything or just using Spectrum. Even Spectrum doesn't always get it right but it is much better then any turbo lora. Thanks for the suggestion any ways.
Its too bad that Fal guy never released his version. Apparently its even better at following prompts and produces better quality then the original model at 10x the speed. I find that very hard to believe but apparently a lot of people who paid to try it confirm this. Just wish he never said he would release and then back out. He must be making a lot of money on it. I just hope the minimax team gets a share of his profits. Pretty damn low if that's not the case and his making bank on a free model.
Agreed, making money off open source is... just greed imo. Donations should be enough.
The lora i use i agree it listens like crap lol and fast hands can be a problem. I have been mostly testing it with fun control as it doesnt need to listen as much to the prompt lol.
My upscale works in batches, you can upscale any video to any res at any batch count you want, its very low vram.
Upscaling does not have the same deterioration as extending. It follows the contrast of the original, not the extend. This is why making 0.2mp 40 seconds works and can be upscaled in batches.
upscaling v2 has a subgraph set that will batch upscale any video you load. You dont need the latents or anything. You can do it one by one or daisy chain the subgraphs to do all batches one at a time.
https://civitai.red/models/2847150/minimax-simple
Thanks very much for sharing your link. Ill try to find some time today or tomorrow to test out your work. By the way, what if I load files that are higher in resolution such as some of my older files that were generated at 0.5. Would I run in to memory problems then and is this why you recommend 0.2 or is 0.2 merely suggested for the speed up?
The size of the input is anything you want. I have upscaled some of my old vace videos that were already nearly 1080p. My vace made videos at about 848x480 and were upscaled via RTX or another method. I needed to crop some of them slightly to fit the size minimax wants but they upscaled fine. This can help remove some of the contrast shift vace produced with batching while upscaling. It clears up faces quite well too.
If you can run a video at 5 seconds at a size, you can upscale any length video in batches of 124 frames at that size. The upscale is 2 steps. Sigma of 0.65, 0.44, 0. Can use sigma 0.72, 0.44, 0. This upscales without changing things. Latent upscales from what i can tell take 3-5 steps and can really change the output. With mine, what you put in, is what comes out, just bigger and cleaner.
I only say 0.2 to 0.4 because minimax can make 40 second videos at that size with my 5060 16gb. I cant make 40 seconds of 0.6 or 1.0. These small outputs can then be upscaled. You can extend a few with the above node instead of making it small.
Very interesting. Thank you for taking the time to explain it.
I was testing out the obvpm latent upscaler with 3 steps and using 0.2 files and was really impressed with the results and the speed. This is until I tried doing multiple files of my extensions. It changed the image and color of all of them so it no longer matched so joining them up for one long video was out of the question.
I thought I would be wasting my time trying any other upscalers thinking they would all lead to the same outcome and would only use an upscaler for just short 20 to 30 second videos without extensions.
If yours can really allow me to run all my extended files through it and I'm able to join them all up to make one long video without it messing up the color,backgrounds and characters, then I really need to give yours a try. I'm just curious why your format, way of doing it isn't more widely known. I can't find any mention of it any where and only the negatives of upscaling for long videos with no known solution.
Darn, I forgot you mentioned you use the RTX Nodes. I could never seem to get that node installed ever. Regardless of what version of Comfy Im using. It does not seem to work with any Portable version of Comfy. I've tried installing manually, through python, and through the comfy manager. It never ever works. Oh well, thanks anyway. If you know of any other workflow that someone offers that does the same as yours but doesn't need the RTX node, I would greatly appreciate any link you could provide. Thanks again for all the info. Appreciate it.
Update.
Someone posted a solution for the Portable version so its working now. I tried the workflow provided on the RTX node Git hub page. It upscaled my 1 minute video to x2 which made an improvement but I find it doesn't look as good as the latent upscaler that's used with the OBVpm workflow. I take it I would get the same quality results through your workflow with the RTX upscaler as I would through there basic workflow. A little disappointed with the RTX node but I guess its better then nothing. At least my curiosity is satisfied now with that particular node.
The rtx node is not needed. you can just resize the video. My flow passes through minimax again just like the latent upscaler. Its not RTX that is doing anything :P. I just added it because it was fast and can offer a decent base to send for upscale. You can resize using any resize node the flow has the second option below the RTX.
The reason to use this is to upscale a video in batches of smaller frame counts to allow larger sizes.
How i use it (working on the full version now)
I generate a 0.2-0.4MP video at the length of 20-35 seconds. I then batch that video into chunks of say 124 or 192 frames. Then resize to 1.4MP with any resize. Then send those batches to render for 2 steps (0.72, 0.44, 0 sigma). Then these batches are merged again at the end.
This allows continuous 30+ seconds at higher res than if you tried to do 30+ at 1.4 or even 2.0 at one time.
Dont get me wrong, i dont think this is better than the obvpm stuff, but the quality is good for simple results.
Example video to video that was done at 0.4mp upscaled to 1.4.
https://civitai.red/images/143654217
Example vace upscale original was 848x480
https://civitai.red/images/143671823
With rtx i needed to do this
RTX install from pythonembeded
python.exe -m pip install -U --no-build-isolation nvidia-vfx --index-url https://pypi.nvidia.com
