Instructions to use Lightricks/LTX-2.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusion Single File
How to use Lightricks/LTX-2.5 with Diffusion Single File:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- LTX-2
How to use Lightricks/LTX-2.5 with LTX-2:
# Install the LTX-2 pipelines git clone https://github.com/Lightricks/LTX-2.git cd LTX-2 uv sync --extra natten
# Download weights from this repo # Substitute filenames from this repo's "Files and versions" if they differ hf download Lightricks/LTX-2.5 \ diffusion_models/<distilled-transformer>.safetensors \ text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ vae/<video-vae>.safetensors \ vae/<audio-vae>.safetensors \ latent_upscale_models/<spatial-upsampler>.safetensors \ latent_upscale_models/<temporal-upsampler>.safetensors \ --local-dir models/LTX-2.5 # DFR requires the detailing IC-LoRA (separate repo; strength is fixed at 0.5) hf download Lightricks/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler --local-dir models/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler# Distilled LTX-2.5 pipeline (fast) uv run python -m ltx_pipelines.distilled \ --transformer-path models/LTX-2.5/diffusion_models/<distilled-transformer>.safetensors \ --text-encoder-path models/LTX-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ --video-vae-path models/LTX-2.5/vae/<video-vae>.safetensors \ --audio-vae-path models/LTX-2.5/vae/<audio-vae>.safetensors \ --spatial-upsampler-path models/LTX-2.5/latent_upscale_models/<spatial-upsampler>.safetensors \ --num-frames 121 \ --prompt "A beautiful sunset over the ocean" \ --output-path output.mp4 # For image-to-video, add: --image path/to/image.jpg 0 0.8# DFR pipeline (higher detail fidelity; optional temporal 2x/4x) uv run python -m ltx_pipelines.dfr_pipeline \ --transformer-path models/LTX-2.5/diffusion_models/<distilled-transformer>.safetensors \ --text-encoder-path models/LTX-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ --video-vae-path models/LTX-2.5/vae/<video-vae>.safetensors \ --audio-vae-path models/LTX-2.5/vae/<audio-vae>.safetensors \ --spatial-upsampler-path models/LTX-2.5/latent_upscale_models/<spatial-upsampler>.safetensors \ --temporal-upsampler-path models/LTX-2.5/latent_upscale_models/<temporal-upsampler>.safetensors \ --detailing-lora models/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler/ltx-2.5-22b-ic-lora-pixel-spatial-upscaler-x2-1.0.safetensors \ --spatial-upscalings 1 \ --temporal-upscalings 1 \ --height 1088 \ --width 1920 \ --num-frames 121 \ --prompt "A beautiful sunset over the ocean" \ --output-path output.mp4 # For 4K: --spatial-upscalings 2 --width 3840 --height 2176 # For image-to-video, add: --image path/to/image.jpg 0 0.8 - Notebooks
- Google Colab
- Kaggle
Regarding the LTX Gemma API Text Encode node...
As an 8GB VRAM user, I'm heavily invested in this node and used it constantly with LTX 2.3. One major pain point, though: if I use a finetune or a quantized version of LTX, the node simply won't let me select it. I have to create a dummy .safetensors file with LTX 2.3 metadata just to trick it. Please keep this in mind for the future - a simple fix would be adding a dropdown in the node to pick the base version (LTX 2, 2.3, 2.5) instead of forcing a manual checkpoint selection.
Also, with LTX 2.5 out, will the node still work? I noticed the tooltip for 'enhance prompt' mentions 'Gemma 3', but LTX 2.5 uses Gemma 4. I know it might just be a label, but it's a bit concerning.
Regardless, huge thanks for making API text encoding possible - it's a genuine lifesaver for low VRAM. I do think your ComfyUI LTXVideo node pack needs a slight update, though (I tested the latest version, which had changes just 5 hours ago). Massive respect for supporting the open-source community.
ive been using the api too but there clearly is something off with it. The audio doesnt work well, prompting a clear vocal text doesnt adhere to it. When i turn off the api and just use local gemma it works fine. The metadata selector still points to 2.3 for the api. So i am guessing something needs to be updated there.
@Abyss35
You are absolutely right. For LTX 2.5 we use our custom fine tuned version of Gemma4 for text encode. Since the text encoder was trained together with the model, using any other text encoder will produce suboptimal results. We are working on adding the API support for Gemma4 as we speak and it should go live very soon.
I will keep you posted
When I use the enhanced prompt with Gemini 4, I get Indian men, some of them even kissing each other. What the hell? But if I use Gemini 3, I get a prompt that is much closer to what I actually want. This is happening with text-to-video. With image-to-video, I get Indian characterization throughout everything that is generated. Is the enhanced prompt not working properly?