8-step 768p V1.0 LoRA released — improved video and audio quality

#48
by lightx2v - opened

We have released the 8-step 768p V1.0 LoRA for MiniMax-H3.

Compared with the previous 4-step version, this release provides improved overall generation quality, with more stable visual details and noticeably better audio quality.

Recommended inference settings

  • Steps: 8
  • Video shift: 6
  • Audio shift: 3
  • Sampler: Euler
  • Scheduler: Simple
  • Resolution: up to 768p

These settings produced the best overall balance of video quality, motion consistency, and audio quality in our testing.

Text to video:

Image to video:

nice job

We need Ref2va update!

still waiting for ref2va

We need Ref2va update!

still waiting for ref2va

did ya'll know you can just... plug the fl2va model into the ref2va h3 node and it functions almost IDENTICALLY to the ref2va model? you can use this lora with it too.

And kijai has the "ref" lora to be used with the FL2V model: https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras

We need Ref2va update!

still waiting for ref2va

did ya'll know you can just... plug the fl2va model into the ref2va h3 node and it functions almost IDENTICALLY to the ref2va model? you can use this lora with it too.

Really? i don't know that! Time to testing

Thanks a lot! can't wait to see the r2fa 8 steps!

Does not work well with other Loras, and at lower screen resolutions(lower than 768p).

Lora strength = 1.0?

please stop waisting time on flf2va, we really need ref2va, thank you!

please stop waisting time on flf2va, we really need ref2va, thank you!

Yeah, i saw merged and fine tune model focusing on flf2v meanwhile the superior model is ref2v

Ref2v is not superior. The image quality is bad compared to fl2va.
That's why we have two models to begin with.

You can use the merged fl2va/ref2va model and use fl2va lora on it

Ref2v is not superior. The image quality is bad compared to fl2va.
That's why we have two models to begin with.

You can use the merged fl2va/ref2va model and use fl2va lora on it

That’s just because you don’t know how to use it. The reference model is what makes it a real productivity tool. I can use it to create an entire movie while keeping the main characters consistent from beginning to end.

Without reference model you can sit and play roulette for whole day hoping to achieve consistency of your character, voice and environment, that's the stupidest shit to wast your time on

How do I say this to you?

Minimax H3 itself has said that the quality of the Ref2VA model is currently lower than FL2VA, and that they’re working on it. Yes, Ref2VA is superior in some aspects, but I hope you realize that FL2VA can also be used as a Ref2VA model. You can use the Hybrid 20–49 model, and its capabilities are actually very close to multi-reference workflows.

So if you run the exact same workflow with Ref2VA and Hybrid FL2VA, Hybrid FL2VA is the winner in my testing.

I’m guessing the LightX2V team already knew about this, while you clearly didn’t.

So before calling someone’s work “stupid shit,” did you actually do any research first?

You don’t need to change your workflow. Just use the Ref2VA workflow and change the model in the Load Diffusion Model node to the Hybrid 20–49 model and use this 8step 1.0 768 lora. That’s it you’re done.

Ref2v needs to be top priority.

I am not calling the work of the team "shit", it is very much appreciated what they do and what they have accomplished, i mentioned that constant re-renderings locally with expectancy to get consistency and accuracy - that is a shit way to spend time.

You don’t need to change your workflow. Just use the Ref2VA workflow and change the model in the Load Diffusion Model node to the Hybrid 20–49 model and use this 8step 1.0 768 lora. That’s it you’re done.

with what settings? strength? scheduler and sampler? audio/video shift?

would really help, thanks.

To stop every discussion, few things to notice to everyone here:

  • Yes, and it's a fact and not contested and prooven. The FL2VA finetuned model is superior, not only on video quality. but also in prompt interpretation and understanding.

  • I think, an very important fact everyone here seems to forgot. BOTH finetuned model can do ALL modes. First there was only one H3 model that do everything, then they decided to separate and finetune two versions from that model to specialize them a little bit. That's why, you have now:
    FL2VA: T2V, I2V and FLF2V
    RE2VA: R2V
    But natively, BOTH models can do T2V, I2V, FLF2V, and R2V, as they come from the main same model.

  • Yes we are all waiting the new RE2VA 4 and 8 step v1.0 768

  • But you know, that you can "easily" do R2V using the FL2VA model. WHY? i just explained above. And in the same time, take advantadge of the far SUPERIOR quality of the FL2VA finetuned model to make R2V, and also the turbo lora fl2va 4 step v1.1 768 amazing quality for RE2VA genration.
    I won't reveal my technique here, as it asked me 3 days of intense testing while developing my custom workflows and nodes, and only my subscribers have this amazig tips, as i implmented it in all my H3 workflows.
    But, just think a little guys, and i'm sure a majority of you will discover this amazing and very simple tips, that is a game changer in term of generation speed/quality for R2V genration.
    Of course, you need some developping skills and to think a lot.
    But i gave you a big clue already above.
    Or just make like everyone, and wait the new turbo lora RE2VA 4 step v1.0 768.

hey @pat11 how can one subscribe you?

hey @pat11 how can one subscribe you?

As i don't found an option to send in PM. I'm adding here:
You can join here:
https://www.patreon.com/cw/Pat3dxcomfyui

To stop every discussion, few things to notice to everyone here:

  • Yes, and it's a fact and not contested and prooven. The FL2VA finetuned model is superior, not only on video quality. but also in prompt interpretation and understanding.

  • I think, an very important fact everyone here seems to forgot. BOTH finetuned model can do ALL modes. First there was only one H3 model that do everything, then they decided to separate and finetune two versions from that model to specialize them a little bit. That's why, you have now:
    FL2VA: T2V, I2V and FLF2V
    RE2VA: R2V
    But natively, BOTH models can do T2V, I2V, FLF2V, and R2V, as they come from the main same model.

  • Yes we are all waiting the new RE2VA 4 and 8 step v1.0 768

  • But you know, that you can "easily" do R2V using the FL2VA model. WHY? i just explained above. And in the same time, take advantadge of the far SUPERIOR quality of the FL2VA finetuned model to make R2V, and also the turbo lora fl2va 4 step v1.1 768 amazing quality for RE2VA genration.
    I won't reveal my technique here, as it asked me 3 days of intense testing while developing my custom workflows and nodes, and only my subscribers have this amazig tips, as i implmented it in all my H3 workflows.
    But, just think a little guys, and i'm sure a majority of you will discover this amazing and very simple tips, that is a game changer in term of generation speed/quality for R2V genration.
    Of course, you need some developping skills and to think a lot.
    But i gave you a big clue already above.
    Or just make like everyone, and wait the new turbo lora RE2VA 4 step v1.0 768.

Yes and no, you're right fundamentally.
But both models differ in the way they work internally and this is why they're two distinct models.

It's acknowledged in the community that if you use FL2VA for pure REF2V you loose the likeness from your reference sheet models.
I use both, have been for hundreds of hours now and there is a difference, it's not blatant but it's noticeable.

REF2V holds the sheets perfectly and your characters look identical, whereas FL2VA drifts off and has less likeliness.

The best we have right now are the hybrid models mixing REF2V and FL2VA, you get a good middle ground of likeness from your sheets, but it's still a tradeoff.

I don't understand people fighting off REF2V like it was the enemy and acting all angry at people asking for it.
Let's be clear here: REF2V is a high-end functionality that, so far, only top of the shelf models like SD and Kling could offer.

The fact an open source model features real REF2V is not a small feat, it's an important milestone and we all get to benefit from it.
So please, put your preferences and ego aside in this scenario.

If you prefer FL2VA for whatever reason, that's totally fine, it's great that the model offers all those possibilities.
But REF2V is the what we need, for obvious reasons, and anyone with experience in training models and generating content using AI know why: consistency is everything, and so far only top of the shelf, expensive models could offer that.

if you cannot wrap your head around this, I'm sorry for you.

To stop every discussion, few things to notice to everyone here:

  • Yes, and it's a fact and not contested and prooven. The FL2VA finetuned model is superior, not only on video quality. but also in prompt interpretation and understanding.

  • I think, an very important fact everyone here seems to forgot. BOTH finetuned model can do ALL modes. First there was only one H3 model that do everything, then they decided to separate and finetune two versions from that model to specialize them a little bit. That's why, you have now:
    FL2VA: T2V, I2V and FLF2V
    RE2VA: R2V
    But natively, BOTH models can do T2V, I2V, FLF2V, and R2V, as they come from the main same model.

  • Yes we are all waiting the new RE2VA 4 and 8 step v1.0 768

  • But you know, that you can "easily" do R2V using the FL2VA model. WHY? i just explained above. And in the same time, take advantadge of the far SUPERIOR quality of the FL2VA finetuned model to make R2V, and also the turbo lora fl2va 4 step v1.1 768 amazing quality for RE2VA genration.
    I won't reveal my technique here, as it asked me 3 days of intense testing while developing my custom workflows and nodes, and only my subscribers have this amazig tips, as i implmented it in all my H3 workflows.
    But, just think a little guys, and i'm sure a majority of you will discover this amazing and very simple tips, that is a game changer in term of generation speed/quality for R2V genration.
    Of course, you need some developping skills and to think a lot.
    But i gave you a big clue already above.
    Or just make like everyone, and wait the new turbo lora RE2VA 4 step v1.0 768.

you must be fun at parties.

To stop every discussion, few things to notice to everyone here:

  • Yes, and it's a fact and not contested and prooven. The FL2VA finetuned model is superior, not only on video quality. but also in prompt interpretation and understanding.

  • I think, an very important fact everyone here seems to forgot. BOTH finetuned model can do ALL modes. First there was only one H3 model that do everything, then they decided to separate and finetune two versions from that model to specialize them a little bit. That's why, you have now:
    FL2VA: T2V, I2V and FLF2V
    RE2VA: R2V
    But natively, BOTH models can do T2V, I2V, FLF2V, and R2V, as they come from the main same model.

  • Yes we are all waiting the new RE2VA 4 and 8 step v1.0 768

  • But you know, that you can "easily" do R2V using the FL2VA model. WHY? i just explained above. And in the same time, take advantadge of the far SUPERIOR quality of the FL2VA finetuned model to make R2V, and also the turbo lora fl2va 4 step v1.1 768 amazing quality for RE2VA genration.
    I won't reveal my technique here, as it asked me 3 days of intense testing while developing my custom workflows and nodes, and only my subscribers have this amazig tips, as i implmented it in all my H3 workflows.
    But, just think a little guys, and i'm sure a majority of you will discover this amazing and very simple tips, that is a game changer in term of generation speed/quality for R2V genration.
    Of course, you need some developping skills and to think a lot.
    But i gave you a big clue already above.
    Or just make like everyone, and wait the new turbo lora RE2VA 4 step v1.0 768.

Yes and no, you're right fundamentally.
But both models differ in the way they work internally and this is why they're two distinct models.

It's acknowledged in the community that if you use FL2VA for pure REF2V you loose the likeness from your reference sheet models.
I use both, have been for hundreds of hours now and there is a difference, it's not blatant but it's noticeable.

REF2V holds the sheets perfectly and your characters look identical, whereas FL2VA drifts off and has less likeliness.

The best we have right now are the hybrid models mixing REF2V and FL2VA, you get a good middle ground of likeness from your sheets, but it's still a tradeoff.

I don't understand people fighting off REF2V like it was the enemy and acting all angry at people asking for it.
Let's be clear here: REF2V is a high-end functionality that, so far, only top of the shelf models like SD and Kling could offer.

The fact an open source model features real REF2V is not a small feat, it's an important milestone and we all get to benefit from it.
So please, put your preferences and ego aside in this scenario.

If you prefer FL2VA for whatever reason, that's totally fine, it's great that the model offers all those possibilities.
But REF2V is the what we need, for obvious reasons, and anyone with experience in training models and generating content using AI know why: consistency is everything, and so far only top of the shelf, expensive models could offer that.

if you cannot wrap your head around this, I'm sorry for you.

You should reread my post. You absolutely don't understand it.
I never said R2V mode was bad.
But that R2V was entirely possible with the FL2VA model.
H3 is a revolution for open source by bringing R2V mode to open source models for the first time.
not only is it a revolution, but it will also enormously boost future models, since all the big companies will want to beat H3, and therefore also integrate the R2V mode.
H3 is the biggest leap forward that we have all been waiting for for years in terms of video generation.
That is a fact and it is not disputed.

I am in no way denigrating R2V mode and its usefulness, which is enormous.
I use almost only the R2V mode, whether with the FL2VA model or natively with the R2V model.

But my post talks about a completely different thing....The fact that we can do R2V with the FL2V model.
With the FL2V model, I can put as a reference image a storyboard, a ref sheet, etc... and H3 will do exactly like with the R2V model, and will generate a completely different scene from the reference image and follow the ref sheet or the 3x3 panel storyboard to make a film.
This is what my post is about. Not that R2V mode is bad and should not be used.

The reason I brought this up is that I see everyone is waiting for the new version of the REF2VA turbo LoRA—one that will deliver results just as good as the FL2V v1.1 turbo LoRA (768) does for the FL2V model at 4 steps.

So, my post explains that it is entirely possible to use R2V mode with the FL2V model, thereby allowing you to use the v1.1 turbo LoRA designed specifically for the FL2V model.
It is a stopgap solution while we wait for a genuine new version of the RE2V 4-step turbo LoRA that matches the performance and quality of the FL2V v1.1 4-step version.

For the record, I don't spend my days training LoRAs; instead, I generate videos—over 50 a day—to develop my custom nodes and ensure the output is perfect based on the parameters.
As a result, I have a thorough understanding of all the new models and how each model and its various modes behave.
I also test every new model to develop new generation tools and integrate them into my image, video, music, and LLM generation applications.
I am no beginner when it comes to AI—quite the opposite.

And believe me, with my technique, character consistency is perfect when using the R2V mode with the FL2V model.
I can't show you any examples, because most of my videos are NSFW for my Patreon subscribers.

Sign up or log in to comment