Image-to-Video
Diffusers
text-to-video
image-text-to-video
video-to-video
text-to-audio-video
image-to-audio-video
image-text-to-audio-video
video-to-audio-video
audio-to-audio-video
audio-video-generation
multimodal
synchronized-audio-video
reference-to-audio-video
Instructions to use rzgar/minimax_h3_fl2va_fp8_e4m3fn with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use rzgar/minimax_h3_fl2va_fp8_e4m3fn with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("rzgar/minimax_h3_fl2va_fp8_e4m3fn", dtype=torch.bfloat16, device_map="cuda") pipe.to("cuda") prompt = "A man with short gray hair plays a red electric guitar." image = load_image( "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png" ) output = pipe(image=image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Notebooks
- Google Colab
- Kaggle
ComfyUI-MiniMaxH3-Text-Enhancer
#6
by K4649 - opened
Can I use this with Sage Attention?
I don't feel the effect of sage attention; perhaps the way the nodes are connected is incorrect.
Those two nodes don’t affect speed, but the internal math. one amplifies the effect of the prompt, and the other mainly affects the look of the male body in T2V scenarios.
together, at 1.2, they make the NSFW-related animations slightly more rhythmic.
It’s mostly useful for people who plan to train NSFW LoRAs.