Instructions to use ntc-ai/minimax-music3-concept-sliders with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ntc-ai/minimax-music3-concept-sliders with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-Music-3", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("ntc-ai/minimax-music3-concept-sliders") prompt = "Same prompt and lyrics. Energy slider Off then Loud +1. Unidirectional lyric-hold LM." image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
MiniMax Music 3 concept sliders
Unidirectional LoRA sliders for MiniMax Music 3.
One scalar adds one musical property while the prompt stays fixed: 0 is off,
+1 is one trained unit of the concept. Opposite poles are separate
LoRAs (Loud and Quiet are two files, not ± on one file).
Current studio catalog is uni-lyric, language-model only. Each shipped
file is weights/<name>-lm-uni-lyric/<name>-lm-uni-lyric_last.safetensors
(unit-normalized, drop in at strength 1), except the four v24 lyric-hold
refreshes (gender-lm-v24-lyrichold, distortion-lm-v24-uni,
breath-lm-v24-uni, joy-lm-v24-uni). Space stays the transformer
exception (space-tf-v6); dust stays dust-lm-v1-faithful-kl-pole03.
Earlier bipolar *-lm-v9 weights remain under weights/ for comparison.
Listen (uni-lyric)
Same prompt, same lyrics, same seed. Only the LoRA scale changes. Player
clips below are uni-lyric at +1 and +2. Full ladders (0 / +1 / +2 plus
caption REFs) live under samples/uni-lyric/.
Energy — Off → Loud
+1 loud
+2 loud
Quiet — Off → Quiet
+1 quiet
+2 quiet
Gender — Off → Female
+1 female
+2 female
Male — Off → Male
+1 male
+2 male
Distortion — Off → Distorted
+1 distorted
+2 distorted
Clean — Off → Clean
+1 clean
+2 clean
Joy — Off → Joy
+1 joy
+2 joy
Somber — Off → Somber
+1 somber
+2 somber
Training code: ntc-ai/sliders-conceptmod
(conceptmod/textsliders/train_lm_slider_music3.py).
Objective (uni-lyric)
Plus-only LoRA on Qwen3Attention. The generation prompt is always the
neutral caption. At LoRA scale s = +1 the prompt-last hidden h(s)
is pulled toward the frozen encode of the + caption h+; at s = 0 it
must reproduce the neutral encode h0. A lyric-token hold keeps the yaml
lyrics span pinned to encode(neu):
Teacher is raw h+ at +1 (last token), h0 at scale 0, and encode(neu)
on the lyrics span only (subscript L). Vocal Details / Global Metadata /
Arrangement are not held. No minus pole, no leftover-gate, no pair-odd. Infer
with the neutral caption + LoRA — not the + caption.
The end term teacher-forces the LoRA'd LM over a frozen base-model composition and penalizes drift of the stop margin:
The slider may move the musical plan; it must not move the stop decision.
Flags:
--lm_target faithful_plus_neu_lyric --pole_mode hidden --endreg_weight 1.0
Opposite concepts ship as separate weights (e.g. energy-lm-uni-lyric
and energy-quiet-lm-uni-lyric). There is no trained −1 on a plus pole.
Catalog
| id | concept | weights |
|---|---|---|
| gender | Female | weights/gender-lm-v24-lyrichold/ |
| male | Male | weights/gender-male-lm-uni-lyric/ |
| energy | Loud | weights/energy-lm-uni-lyric/ |
| quiet | Quiet | weights/energy-quiet-lm-uni-lyric/ |
| distortion | Distorted | weights/distortion-lm-v24-uni/ |
| clean | Clean | weights/distortion-clean-lm-uni-lyric/ |
| tempo | Fast | weights/tempo-lm-uni-lyric/ |
| slow | Slow | weights/tempo-slow-lm-uni-lyric/ |
| live | Live | weights/live-lm-uni-lyric/ |
| studio | Studio | weights/live-studio-lm-uni-lyric/ |
| breath | Breathy | weights/breath-lm-v24-uni/ |
| rapslow | Rap | weights/rapslow-lm-uni-lyric/ |
| triphop | Trip-hop | weights/triphop-lm-uni-lyric/ |
| pop | Pop | weights/triphop-pop-lm-uni-lyric/ |
| rhyme | Rhyme | weights/rhyme-lm-uni-lyric/ |
| prose | Prose | weights/rhyme-prose-lm-uni-lyric/ |
| sexy | Sexy | weights/sexy-lm-uni-lyric/ |
| plain | Plain | weights/sexy-plain-lm-uni-lyric/ |
| tender | Tender | weights/tender-lm-uni-lyric/ |
| fierce | Fierce | weights/tender-fierce-lm-uni-lyric/ |
| grit | Grit | weights/grit-lm-uni-lyric/ |
| smooth | Smooth | weights/grit-smooth-lm-uni-lyric/ |
| joy | Joy | weights/joy-lm-v24-uni/ |
| somber | Somber | weights/joy-somber-lm-uni-lyric/ |
| yearn | Yearning | weights/yearn-lm-uni-lyric/ |
| settled | Settled | weights/yearn-settled-lm-uni-lyric/ |
| hurt | Hurt | weights/hurt-lm-uni-lyric/ |
| numb | Numb | weights/hurt-numb-lm-uni-lyric/ |
| dust | Tape | weights/dust-lm-v1-faithful-kl-pole03/ |
| space | Wet | weights/space-tf-v6/ (transformer) |
Play samples/uni-lyric/<name>-lm-uni-lyric/ in filename order (0, +1, +2,
then caption REFs).
ComfyUI
Each shipped file has a sibling *_comfyui.safetensors (LM files use
text_encoders.model.layers.*.self_attn.*). Drop the Comfy file in
ComfyUI/models/loras/ and load it with Load LoRA. Strength is the slider
scale (0 off, +1 the baked unit).
ComfyUI support is untested.
Which stage a slider attaches to
MiniMax Music 3 generates in two stages:
- A Qwen3 language model writes a plan — who is singing, melody, rhythm, arrangement.
- A flow transformer renders that plan — timbre, loudness, tone, space.
uni-lyric attaches to stage 1. Space is the catalog exception on stage 2 (reverb is a rendering property).
- Downloads last month
- -