becausereasons commited on
Commit
23dab68
·
verified ·
1 Parent(s): 8205691

New-trainer generation: chnsn_montmartre (step 300) + chnsn_cabaret (step 350), 6 demos, 2 prompts

Browse files
.gitattributes CHANGED
@@ -38,3 +38,9 @@ demos/chnsn_rive_gauche__acoustic_narrative.mp3 filter=lfs diff=lfs merge=lfs -t
38
  demos/chnsn_rive_gauche__duet_letter.mp3 filter=lfs diff=lfs merge=lfs -text
39
  demos/chnsn_rive_gauche__yeye_baroque.mp3 filter=lfs diff=lfs merge=lfs -text
40
  demos/chnsn_rive_gauche__female_dark_waltz.mp3 filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
38
  demos/chnsn_rive_gauche__duet_letter.mp3 filter=lfs diff=lfs merge=lfs -text
39
  demos/chnsn_rive_gauche__yeye_baroque.mp3 filter=lfs diff=lfs merge=lfs -text
40
  demos/chnsn_rive_gauche__female_dark_waltz.mp3 filter=lfs diff=lfs merge=lfs -text
41
+ demos/chnsn_cabaret__swing_bigband.mp3 filter=lfs diff=lfs merge=lfs -text
42
+ demos/chnsn_cabaret__yeye_baroque_clip05.mp3 filter=lfs diff=lfs merge=lfs -text
43
+ demos/chnsn_montmartre__female_dark_waltz_clip05.mp3 filter=lfs diff=lfs merge=lfs -text
44
+ demos/chnsn_montmartre__female_torch.mp3 filter=lfs diff=lfs merge=lfs -text
45
+ demos/chnsn_montmartre__male_dark_crescendo.mp3 filter=lfs diff=lfs merge=lfs -text
46
+ demos/chnsn_montmartre__male_piano_live.mp3 filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -1,240 +1,294 @@
1
- ---
2
- license: cc-by-nc-4.0
3
- base_model:
4
- - m-a-p/YuE2-3B
5
- - Comfy-Org/YuE2
6
- pipeline_tag: text-to-audio
7
- language:
8
- - fr
9
- tags:
10
- - lora
11
- - yue2
12
- - music-generation
13
- - text-to-music
14
- - song-generation
15
- - french
16
- - chanson
17
- - chanson-francaise
18
- - orchestral-pop
19
- - ye-ye
20
- - spoken-word
21
- - comfyui
22
- - fs_audio
23
- ---
24
-
25
- # CHNSN — Chanson Française LoRAs for YuE2
26
-
27
- Two checkpoints of one artist-style LoRA that push [YuE2-3B](https://huggingface.co/m-a-p/YuE2-3B) into **classic French chanson**: theatrical baritone and soprano leads with real vibrato, sweeping strings, accordion, piano, brass swells, upright bass, brushed drums, and the mid-century orchestral sound behind them. Vocals are sung in French. Both male and female leads are in the data, and the prompt decides which one you get. It also does something the base model is not built for: a **spoken passage** inside a song (see *A read letter* below).
28
-
29
- Each file patches **both halves** of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for both: **`chnsn`**.
30
-
31
- | | LoRA | File | Character | Start here |
32
- |---|---|---|---|---|
33
- | 🎻 | **CHNSN Rive Gauche** | `chnsn_rive_gauche.safetensors` | Step 200, the decoder-loss minimum. The more supple of the two: acoustic narrative, yé-yé, female leads, waltz meters, and the sung-letter duet all came from this file. | clip 1.0 / model 1.0, cfg 1.0 |
34
- | 🎺 | **CHNSN Grand Boulevard** | `chnsn_grand_boulevard.safetensors` | Step 350, the final checkpoint. Tighter, more orchestral, writes shorter chanson-length songs. The orchestral ballad demo. | clip 1.0 / model 1.0, cfg 1.0 |
35
-
36
- ## Listen
37
-
38
- All demos use an original French lyric (nine tagged sections, a Paris-in-the-rain chanson), seed 7, baseline recipe, 32 steps `dpm_2` / `sgm_uniform`, 360 s cap, no post-processing.
39
-
40
- **CHNSN Rive Gauche — a read letter** (duet), prompt `prompts/duet_spoken_letter.txt`. Female lead on the verses and chorus, then a male baritone takes a letter written as prose and delivers it as a declaimed verse, then the chorus returns. 6/8, 90 BPM, D minor, 3:06:
41
-
42
- <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_rive_gauche__duet_letter.mp3"></audio>
43
-
44
- **CHNSN Rive Gauche — yé-yé baroque**, prompt `prompts/female_yeye_baroque.txt`: youthful female lead, harpsichord, fuzz bass, mellotron flutes, brass stabs (125 BPM, C major, 4:09):
45
-
46
- <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_rive_gauche__yeye_baroque.mp3"></audio>
47
-
48
- **CHNSN Rive Gauche — female dark waltz**, prompt `prompts/female_dark_waltz.txt`: a dramatic female lead with vibrato and declaimed verses over a dark slow waltz, accordion, dissonant strings and tape textures (6/8, 90 BPM, D minor, 3:18, 83 % of vocal bars sung):
49
-
50
- <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_rive_gauche__female_dark_waltz.mp3"></audio>
51
-
52
- **CHNSN Rive Gauche — acoustic narrative**, prompt `prompts/acoustic_narrative.txt`: warm baritone storytelling over nylon-string guitar, upright bass and soft strings (105 BPM, D minor, 3:44):
53
-
54
- <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_rive_gauche__acoustic_narrative.mp3"></audio>
55
-
56
- **CHNSN Grand Boulevard — orchestral ballad**, prompt `prompts/orchestral_ballad.txt`: dramatic baritone with vibrato, sweeping strings, piano, brass swells (86 BPM, D minor, 3:26):
57
-
58
- <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_grand_boulevard__orchestral_ballad.mp3"></audio>
59
-
60
- ## Quick start (ComfyUI)
61
-
62
- The LoRAs are in the fused-key layout that Comfy's YuE2 implementation uses (`text_encoders.*` for the planner, `diffusion_model.*` for the decoder). They were trained with, and load through, the **[FS_Audio Suite](https://github.com/KytraScript/ComfyUI-FS_Audio_Suite)** node pack.
63
-
64
- 1. ComfyUI ≥ v0.36.0 (native YuE2 support) and the FS_Audio Suite custom node pack.
65
- 2. Base model: `yue2_3b_bf16.safetensors` from [Comfy-Org/YuE2](https://huggingface.co/Comfy-Org/YuE2) in `models/checkpoints/`.
66
- 3. Drop one `chnsn_*.safetensors` into `models/loras/`.
67
- 4. Chain the nodes:
68
-
69
- ```
70
- 🧩 FS_Audio Lora Loader ──loras──▶ 🎤 FS_Audio Model Loader ──pipe──▶ 🎵 FS_Audio Sampler ──▶ 💿 FS_Audio Output
71
- lora_name = chnsn_rive_gauche.safetensors yue2_checkpoint = yue2_3b_bf16 style = <prompt, starts with "chnsn,">
72
- strength_clip = 1.0 (planner) melody_transcriber = none lyrics = <tagged French lyric blocks>
73
- strength_model = 1.0 (decoder) score_mode = full
74
- ```
75
-
76
- Sampler settings used for every demo:
77
-
78
- | widget | value |
79
- |---|---|
80
- | steps / sampler / scheduler | 32 / `dpm_2` / `sgm_uniform` |
81
- | score_mode | `full` |
82
- | song_length_cap | 360 (chanson-length lyrics plan to 3–5 minutes; the cap rarely triggers) |
83
- | repetition_penalty | 1.2 |
84
-
85
- Two recipes:
86
-
87
- | | Baseline (all demos) | Wild |
88
- |---|---|---|
89
- | strength_clip (planner) | 1.0 | 1.0 |
90
- | strength_model (decoder) | 1.0 | 1.0 |
91
- | Weirdness (cfg) | 1.0 | **1.4** |
92
- | score_temperature | 0.7 | **0.9** |
93
- | music_temperature | 1.0 | **1.2** |
94
-
95
- Headless, the same graph as an API prompt:
96
-
97
- ```python
98
- prompt = {
99
- "1": {"class_type": "FSAudioLoraLoader", "inputs": {"lora_name": "chnsn_rive_gauche.safetensors", "strength_model": 1.0, "strength_clip": 1.0}},
100
- "2": {"class_type": "FSAudioModelLoader", "inputs": {"yue2_checkpoint": "yue2_3b_bf16.safetensors", "melody_transcriber": "none", "loras": ["1", 0]}},
101
- "3": {"class_type": "FSAudioSampler", "inputs": {"pipe": ["2", 0], "style": STYLE, "lyrics": LYRICS, "seed": 7,
102
- "song_length_cap": 360, "score_mode": "full", "steps": 32, "sampler": "dpm_2", "scheduler": "sgm_uniform",
103
- "Weirdness (cfg)": 1.0, "score_temperature": 0.7, "music_temperature": 1.0, "repetition_penalty": 1.2}},
104
- "4": {"class_type": "FSAudioOutput", "inputs": {"song": ["3", 0], "score": ["3", 1], "info": ["3", 2], "filename_prefix": "chnsn"}},
105
- }
106
- ```
107
-
108
- ## Prompting
109
-
110
- ### Style prompt
111
-
112
- Start with the trigger, then write **one descriptive sentence** in this order: language → genre → vocal → instruments → mood → production → BPM. This is the format the planner was trained on. Bare tag lists produce odd plans.
113
-
114
- ```
115
- chnsn, French, classic French chanson, dramatic expressive male baritone vocal with rich vibrato and theatrical phrasing, sweeping orchestral strings, acoustic piano, brass swells, upright bass, soft brushed drums, poignant, romantic and deeply melancholic, grand cinematic arrangement with emotional crescendos, 81 BPM
116
- ```
117
-
118
- Ready-made prompts in `prompts/`:
119
-
120
- | file | what it gets you |
121
- |---|---|
122
- | `orchestral_ballad.txt` | the Grand Boulevard demo: dramatic baritone, strings, piano, brass |
123
- | `acoustic_narrative.txt` | warm baritone storytelling, nylon guitar, upright bass, soft strings |
124
- | `swing_bigband.txt` | 1960s big-band swing chanson; rendered but the planner leaves long instrumental stretches (only a third of vocal bars sung), so treat it as an instrumental-leaning prompt |
125
- | `female_yeye_baroque.txt` | the yé-yé demo: youthful female lead, harpsichord, fuzz bass, mellotron |
126
- | `female_dark_waltz.txt` | dramatic female with vibrato over a dark 6/8 waltz, accordion, dissonant strings, tape loops (demo) |
127
- | `female_torch_minimal.txt` | smoky low-register female torch song, felt piano, glass harmonica, long silences; use the wild recipe (rendered, not in the demos) |
128
- | `duet_spoken_letter.txt` | the duet demo: female lead plus a male baritone narrator reading a letter |
129
-
130
- "Female vocal" in the prompt reliably gives a female lead even though only three of the fifteen training songs have one. Waltz words ("slow waltz", "dark waltz") put the planner in 6/8; "yé-yé", "harpsichord", "fuzz bass" push it to 125 BPM in a major key.
131
-
132
- ### A read letter (spoken passages)
133
-
134
- YuE2 has no spoken-word mode: the planner writes a pitched melody for every vocal bar. But two things together get a passage *delivered* rather than sung:
135
-
136
- 1. Write the passage as **prose**: long lines, no rhyme, no repetition, second person, like a letter. Put it in its own section.
137
- 2. In the style prompt, name the narrator and say what he does: `a deep low male baritone narrator who does not sing but reads a letter aloud in a calm monotone spoken voice over the bridge, recited parlé passage with no melody ... the spoken passage close-miked and dry`.
138
-
139
- The section tag then picks the flavour. Under a normal **`[Bridge]`** tag the planner writes the letter as a declaimed verse on a simple melody and hands it to the male voice: that is the duet demo. Under an unfamiliar **`[Spoken]`** tag the planner writes that stretch as an *instrumental interlude with no vocal melody at all*, and the decoder still voices every word over the band, at about 2.5 words per second: proper recitation. The `[Spoken]` render carried all of the letter; the `[Bridge]` render sang about half of it before returning to the chorus. Check the `.abc` next to the render: a `% interlude` with rest-only `V: Vocal` bars at that spot means it worked.
140
-
141
- ### Lyrics
142
-
143
- Tagged French blocks. Use `[Intro]`, `[Verse 1]`, `[Verse 2]`, `[Chorus]`, `[Bridge]`, `[Outro]` (and `[Spoken]` for narration); number the verses and write a repeated chorus out again where you want it sung. Standard orthography with accents and apostrophes. Shape that made the demos:
144
-
145
- ```
146
- [Intro]
147
- two short lines, the image the song returns to
148
-
149
- [Verse 1]
150
- four lines, ten to twelve syllables, rhymed in pairs
151
-
152
- [Verse 2]
153
- four more lines
154
-
155
- [Chorus]
156
- four lines, the hook phrase at the top of the first two
157
-
158
- [Verse 3]
159
- four lines
160
-
161
- [Chorus]
162
- (repeat, written out)
163
-
164
- [Bridge]
165
- four lines, or a prose letter for the narrator
166
-
167
- [Chorus]
168
- (repeat, written out)
169
-
170
- [Outro]
171
- the intro image again, trailing off
172
- ```
173
-
174
- The planner sometimes drops the bridge on the orchestral prompt and goes verse–chorus–verse–chorus–outro; the acoustic and swing prompts keep it. Chanson-length lyrics like this plan to 3–5 minutes.
175
-
176
- ### Which knob does what
177
-
178
- - **Checkpoint** is the first choice: Rive Gauche (step 200) for anything acoustic, female, experimental or spoken; Grand Boulevard (step 350) for the big orchestral ballad.
179
- - **strength_model (decoder)** 1.0 → 1.5 makes the voice rougher and more present without touching the writing. **strength_clip (planner) above 1.0 collapses the vocal** on YuE2 artist LoRAs in general; keep it at 1.0.
180
- - **Weirdness (cfg)** changes only the sound. **music_temperature / score_temperature** change the writing; 1.2 / 0.9 is the wild recipe used for the torch-song prompt.
181
- - **Seeds** decide length as much as the cap does; the same lyric planned 3:26 on one checkpoint and 4:28 on the other.
182
-
183
- ## Training
184
-
185
- Trained on a single RTX 5090 with the FS_Audio Suite **Artist Trainer** on Comfy's own YuE2 weights. One run trains the planner LoRA and the decoder LoRA together; the file also carries full `vae2llm` / `llm2vae` projection diffs. Both published files come from the same run.
186
-
187
- | | Rive Gauche | Grand Boulevard |
188
- |---|---|---|
189
- | dataset | 15 songs, 45 minutes: French chanson from the 1950s to the 1980s, mostly orchestral male-baritone chanson with a few acoustic guitar narratives and three female-led songs | same |
190
- | checkpoint | **step 200** | step 350 (final) |
191
- | planner / decoder steps | 350 / 600 | 350 / 600 |
192
- | artist (planner) loss | 5.220 → 4.606 | 5.220 → 4.582 |
193
- | decoder loss at checkpoint | **1.148** (the minimum) | 1.162 |
194
- | wall time | 21 minutes for the whole run, dataset build included | |
195
-
196
- Hyper-parameters: rank 64 planner / rank 32 decoder, LR 3e-5 planner / 4e-5 decoder / 2e-5 I/O projections, artist fraction 0.5 vs regularizer pack, KL 0.1, batch 2 songs, 8192 max tokens, 750-frame windows, EMA 0.99, one song held out, score-first 0, no transcription.
197
-
198
- **Why two checkpoints.** The trainer names its `_best` file by planner loss, which keeps falling to the end. The decoder loss bottomed at step 200 (1.148) and drifted back up to 1.162 by step 350 on this 15-song set, a mild version of the overfit seen on other small sets. Step 200 is the pick by the decoder curve, and three of the four demos are from it. The step-350 file is published because the orchestral ballad rendered from it was preferred by ear, and the drift is small enough that both are usable.
199
-
200
- ### Captioning your own dataset
201
-
202
- The planner learns from the captions as much as from the audio, so the caption format decides whether your prompts work later. What worked here:
203
-
204
- - **One descriptive sentence per song, trigger first**, in the order you will prompt in: language → genre → vocal → instruments → mood → production → BPM. Decade words ("1960s French chanson") are useful descriptors.
205
- - **No tag lists, no section scaffolding** in the caption.
206
- - **Measure BPM, don't guess**, and sanity-check it: beat trackers double slow ballads and halve fast ones (a lively yé-yé track came out as 72 and is 144; a piano ballad came out as 136 and is 68). Compare the number with the mood words in the same caption.
207
- - A vision-language model can draft the captions from the audio, but **check the vocal gender by hand**: a high male tenor got labelled "female lead" here. The three real female leads were labelled correctly, so this is a grep-and-listen pass, not a blind find-and-replace.
208
- - **Lyrics as tagged blocks** with standard orthography, accents and apostrophes. Transcribe at high confidence; a mis-heard lyric teaches the planner the wrong syllable count for the bar.
209
- - **Audio prep**: FLAC 44.1 kHz / 16-bit, ≤ 320 s. MP3 rips with a damaged leading frame crash the dataset builder ("Header missing"); transcode with `ffmpeg -err_detect ignore_err`.
210
- - **Trainer flags**: `score-first 0` and no automatic transcription.
211
-
212
- ## Known limitations
213
-
214
- - French vocals only.
215
- - Female leads come from three training songs; they render well on the prompts above but the range of female timbres is narrower than the male ones.
216
- - The big-band swing prompt leaves most bars instrumental. Add "sung throughout" and a denser lyric, or change the seed.
217
- - The planner may drop a `[Bridge]` on orchestral prompts. If the bridge matters, try the acoustic prompt family or the `[Spoken]` tag.
218
- - No instrumental-only mode is baked in; the LoRAs assume a lyric.
219
-
220
- ## Files
221
-
222
- ```
223
- chnsn_rive_gauche.safetensors 177 MB step 200, planner + decoder LoRA (bf16 weights, fp32 projection diffs)
224
- chnsn_grand_boulevard.safetensors 177 MB step 350, same layout
225
- demos/ mp3 renders (192 kbps from the FLAC masters), seed 7 throughout
226
- prompts/ style prompts
227
- ```
228
-
229
- ## Support
230
-
231
- These LoRAs are trained on my own GPU and released free. If they're useful to you and you'd like to chip in for compute, there's a Ko-fi: **[ko-fi.com/becausereasons](https://ko-fi.com/becausereasons)** <3
232
-
233
- ## License and credits
234
-
235
- Weights are released under **CC BY-NC 4.0**, inherited from the YuE2-3B base model. Non-commercial use only; attribute "CHNSN LoRAs by becausereasons".
236
-
237
- - [YuE2](https://huggingface.co/m-a-p/YuE2-3B) by the Multimodal Art Projection (m-a-p) team; ComfyUI repack by [Comfy-Org](https://huggingface.co/Comfy-Org/YuE2).
238
- - [ComfyUI-FS_Audio_Suite](https://github.com/KytraScript/ComfyUI-FS_Audio_Suite) by KytraScript / The Fixed Seed Company: inference nodes and the artist trainer.
239
- - Sister releases: [CNZN canzone italiana](https://huggingface.co/becausereasons/yue2-cnzn-canzone-italiana), [MLTNT militant roots reggae](https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae).
240
- - Trained and documented by becausereasons, September 2026.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-nc-4.0
3
+ base_model:
4
+ - m-a-p/YuE2-3B
5
+ - Comfy-Org/YuE2
6
+ pipeline_tag: text-to-audio
7
+ language:
8
+ - fr
9
+ tags:
10
+ - lora
11
+ - yue2
12
+ - music-generation
13
+ - text-to-music
14
+ - song-generation
15
+ - french
16
+ - chanson
17
+ - chanson-francaise
18
+ - orchestral-pop
19
+ - ye-ye
20
+ - spoken-word
21
+ - comfyui
22
+ - fs_audio
23
+ - ai-toolkit
24
+ ---
25
+
26
+ # CHNSN — Chanson Française LoRAs for YuE2
27
+
28
+ Artist-style LoRAs that push [YuE2-3B](https://huggingface.co/m-a-p/YuE2-3B) into **classic French chanson**: theatrical baritone and soprano leads with real vibrato, sweeping strings, accordion, piano, brass swells, upright bass, brushed drums, and the mid-century orchestral sound behind them. Vocals are sung in French. Both male and female leads are in the data, and the prompt decides which one you get. It also does something the base model is not built for: a **spoken passage** inside a song (see *A read letter* below).
29
+
30
+ Each file patches **both halves** of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all of them: **`chnsn`**.
31
+
32
+ **September 2026 update: a new-trainer generation.** `chnsn_montmartre` and `chnsn_cabaret` were trained on the same 15 songs with a different trainer (the experimental YuE2 support in [Ostris AI Toolkit](https://github.com/ostris/ai-toolkit)). The voices are more expressive and more idiomatic than on the two earlier files, male and female alike. They are early checkpoints on purpose: on a set this small the later ones pick up an audible high-end grit, explained under *Known limitations*. The two earlier files stay available and unchanged.
33
+
34
+ *Last updated: 20 September 2026.*
35
+
36
+ | | LoRA | File | Character | Start here |
37
+ |---|---|---|---|---|
38
+ | 🌙 | **CHNSN Montmartre** | `chnsn_montmartre.safetensors` | New-trainer generation, step 300. The most expressive voices of the set, clean breaths and whispers, wide range across the prompts below. | clip 1.0 / model 1.0; drop to **clip 0.5** if a song will not end |
39
+ | 🎭 | **CHNSN Cabaret** | `chnsn_cabaret.safetensors` | New-trainer generation, step 350. A touch more settled and more orchestral than Montmartre; the swing and yé-yé demos. Very occasionally a trace of high-end grit on breaths. | clip 1.0 or **clip 0.5**, model 1.0 |
40
+ | 🎻 | **CHNSN Rive Gauche** | `chnsn_rive_gauche.safetensors` | Step 200, the decoder-loss minimum. The more supple of the two: acoustic narrative, yé-yé, female leads, waltz meters, and the sung-letter duet all came from this file. | clip 1.0 / model 1.0, cfg 1.0 |
41
+ | 🎺 | **CHNSN Grand Boulevard** | `chnsn_grand_boulevard.safetensors` | Step 350, the final checkpoint. Tighter, more orchestral, writes shorter chanson-length songs. The orchestral ballad demo. | clip 1.0 / model 1.0, cfg 1.0 |
42
+
43
+ ## Listen: new-trainer generation
44
+
45
+ Same original French lyric, seed 7, 32 steps `dpm_2` / `sgm_uniform`, 360 s cap, no post-processing. `clip` is `strength_clip` (the planner), `model` is `strength_model` (the decoder).
46
+
47
+ **Montmartre, female torch song** — `prompts/female_torch_minimal.txt`, full strength (clip 1.0 / model 1.0). A smoky low female lead with whispered passages over felt piano and bowed bass; the whispers stay clean (6/8, 96 BPM, D minor, 2:51):
48
+
49
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_montmartre__female_torch.mp3"></audio>
50
+
51
+ **Montmartre, male dark crescendo** — `prompts/male_dark_crescendo.txt`, full strength. A dark baritone with wide vibrato, hushed verses rising to a full-voiced orchestral climax (6/8, 90 BPM, D minor, 3:35):
52
+
53
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_montmartre__male_dark_crescendo.mp3"></audio>
54
+
55
+ **Montmartre, male piano ballad, live** — `prompts/male_piano_live.txt`, full strength. A resonant baritone with heavy vibrato, hushed verses and a belted chorus over grand piano (70 BPM, D minor, 3:34):
56
+
57
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_montmartre__male_piano_live.mp3"></audio>
58
+
59
+ **Montmartre, female dark waltz** — `prompts/female_dark_waltz.txt`, clip 0.5 / model 1.0 (6/8, 90 BPM, D minor, 3:24):
60
+
61
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_montmartre__female_dark_waltz_clip05.mp3"></audio>
62
+
63
+ **Cabaret, big-band swing** — `prompts/swing_bigband.txt`, full strength (115 BPM, F major, 3:26):
64
+
65
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_cabaret__swing_bigband.mp3"></audio>
66
+
67
+ **Cabaret, yé-yé baroque** — `prompts/female_yeye_baroque.txt`, clip 0.5 / model 1.0 (125 BPM, C major, 4:12):
68
+
69
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_cabaret__yeye_baroque_clip05.mp3"></audio>
70
+
71
+ ### Using the new-trainer files
72
+
73
+ They load exactly like the others (same fused-key layout, `text_encoders.*` planner and `diffusion_model.*` decoder; the LoRA matrices are stored as `lora_A` / `lora_B`, and there are no projection diffs), through the same FS_Audio graph shown in *Quick start*.
74
+
75
+ - **Start at clip 1.0 / model 1.0.** At steps 300 and 350 full strength is safe on every prompt family on this page.
76
+ - **If the planner writes a song that will not end** (a score stuck in one section, music with no singing), set clip to 0.5 or 0.8. In our tests every runaway score at clip 1.0 was healthy again at 0.8.
77
+ - **Use a 360 s cap.** In score mode the song length is decided up front by the score, so the cap only truncates.
78
+
79
+ What the new trainer does differently, what it costs, and how to spot a runaway score from the `.abc` sidecar are written up on the sister page: [CNZN, *Two trainers: what we learned*](https://huggingface.co/becausereasons/yue2-cnzn-canzone-italiana#two-trainers-what-we-learned).
80
+
81
+ ## Listen: the two earlier files
82
+
83
+ All demos use an original French lyric (nine tagged sections, a Paris-in-the-rain chanson), seed 7, baseline recipe, 32 steps `dpm_2` / `sgm_uniform`, 360 s cap, no post-processing.
84
+
85
+ **CHNSN Rive Gauche — a read letter** (duet), prompt `prompts/duet_spoken_letter.txt`. Female lead on the verses and chorus, then a male baritone takes a letter written as prose and delivers it as a declaimed verse, then the chorus returns. 6/8, 90 BPM, D minor, 3:06:
86
+
87
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_rive_gauche__duet_letter.mp3"></audio>
88
+
89
+ **CHNSN Rive Gauche — yé-yé baroque**, prompt `prompts/female_yeye_baroque.txt`: youthful female lead, harpsichord, fuzz bass, mellotron flutes, brass stabs (125 BPM, C major, 4:09):
90
+
91
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_rive_gauche__yeye_baroque.mp3"></audio>
92
+
93
+ **CHNSN Rive Gauche — female dark waltz**, prompt `prompts/female_dark_waltz.txt`: a dramatic female lead with vibrato and declaimed verses over a dark slow waltz, accordion, dissonant strings and tape textures (6/8, 90 BPM, D minor, 3:18, 83 % of vocal bars sung):
94
+
95
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_rive_gauche__female_dark_waltz.mp3"></audio>
96
+
97
+ **CHNSN Rive Gauche — acoustic narrative**, prompt `prompts/acoustic_narrative.txt`: warm baritone storytelling over nylon-string guitar, upright bass and soft strings (105 BPM, D minor, 3:44):
98
+
99
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_rive_gauche__acoustic_narrative.mp3"></audio>
100
+
101
+ **CHNSN Grand Boulevard — orchestral ballad**, prompt `prompts/orchestral_ballad.txt`: dramatic baritone with vibrato, sweeping strings, piano, brass swells (86 BPM, D minor, 3:26):
102
+
103
+ <audio controls src="https://huggingface.co/becausereasons/yue2-chnsn-chanson-francaise/resolve/main/demos/chnsn_grand_boulevard__orchestral_ballad.mp3"></audio>
104
+
105
+ ## Quick start (ComfyUI)
106
+
107
+ The LoRAs are in the fused-key layout that Comfy's YuE2 implementation uses (`text_encoders.*` for the planner, `diffusion_model.*` for the decoder). They were trained with, and load through, the **[FS_Audio Suite](https://github.com/KytraScript/ComfyUI-FS_Audio_Suite)** node pack.
108
+
109
+ 1. ComfyUI ≥ v0.36.0 (native YuE2 support) and the FS_Audio Suite custom node pack.
110
+ 2. Base model: `yue2_3b_bf16.safetensors` from [Comfy-Org/YuE2](https://huggingface.co/Comfy-Org/YuE2) in `models/checkpoints/`.
111
+ 3. Drop one `chnsn_*.safetensors` into `models/loras/`.
112
+ 4. Chain the nodes:
113
+
114
+ ```
115
+ 🧩 FS_Audio Lora Loader ──loras──▶ 🎤 FS_Audio Model Loader ──pipe──▶ 🎵 FS_Audio Sampler ──▶ 💿 FS_Audio Output
116
+ lora_name = chnsn_rive_gauche.safetensors yue2_checkpoint = yue2_3b_bf16 style = <prompt, starts with "chnsn,">
117
+ strength_clip = 1.0 (planner) melody_transcriber = none lyrics = <tagged French lyric blocks>
118
+ strength_model = 1.0 (decoder) score_mode = full
119
+ ```
120
+
121
+ Sampler settings used for every demo:
122
+
123
+ | widget | value |
124
+ |---|---|
125
+ | steps / sampler / scheduler | 32 / `dpm_2` / `sgm_uniform` |
126
+ | score_mode | `full` |
127
+ | song_length_cap | 360 (chanson-length lyrics plan to 3–5 minutes; the cap rarely triggers) |
128
+ | repetition_penalty | 1.2 |
129
+
130
+ Two recipes:
131
+
132
+ | | Baseline (all demos) | Wild |
133
+ |---|---|---|
134
+ | strength_clip (planner) | 1.0 | 1.0 |
135
+ | strength_model (decoder) | 1.0 | 1.0 |
136
+ | Weirdness (cfg) | 1.0 | **1.4** |
137
+ | score_temperature | 0.7 | **0.9** |
138
+ | music_temperature | 1.0 | **1.2** |
139
+
140
+ Headless, the same graph as an API prompt:
141
+
142
+ ```python
143
+ prompt = {
144
+ "1": {"class_type": "FSAudioLoraLoader", "inputs": {"lora_name": "chnsn_rive_gauche.safetensors", "strength_model": 1.0, "strength_clip": 1.0}},
145
+ "2": {"class_type": "FSAudioModelLoader", "inputs": {"yue2_checkpoint": "yue2_3b_bf16.safetensors", "melody_transcriber": "none", "loras": ["1", 0]}},
146
+ "3": {"class_type": "FSAudioSampler", "inputs": {"pipe": ["2", 0], "style": STYLE, "lyrics": LYRICS, "seed": 7,
147
+ "song_length_cap": 360, "score_mode": "full", "steps": 32, "sampler": "dpm_2", "scheduler": "sgm_uniform",
148
+ "Weirdness (cfg)": 1.0, "score_temperature": 0.7, "music_temperature": 1.0, "repetition_penalty": 1.2}},
149
+ "4": {"class_type": "FSAudioOutput", "inputs": {"song": ["3", 0], "score": ["3", 1], "info": ["3", 2], "filename_prefix": "chnsn"}},
150
+ }
151
+ ```
152
+
153
+ ## Prompting
154
+
155
+ ### Style prompt
156
+
157
+ Start with the trigger, then write **one descriptive sentence** in this order: language → genre → vocal → instruments → mood → production → BPM. This is the format the planner was trained on. Bare tag lists produce odd plans.
158
+
159
+ ```
160
+ chnsn, French, classic French chanson, dramatic expressive male baritone vocal with rich vibrato and theatrical phrasing, sweeping orchestral strings, acoustic piano, brass swells, upright bass, soft brushed drums, poignant, romantic and deeply melancholic, grand cinematic arrangement with emotional crescendos, 81 BPM
161
+ ```
162
+
163
+ Ready-made prompts in `prompts/`:
164
+
165
+ | file | what it gets you |
166
+ |---|---|
167
+ | `orchestral_ballad.txt` | the Grand Boulevard demo: dramatic baritone, strings, piano, brass |
168
+ | `acoustic_narrative.txt` | warm baritone storytelling, nylon guitar, upright bass, soft strings |
169
+ | `swing_bigband.txt` | 1960s big-band swing chanson; rendered but the planner leaves long instrumental stretches (only a third of vocal bars sung), so treat it as an instrumental-leaning prompt |
170
+ | `female_yeye_baroque.txt` | the yé-yé demo: youthful female lead, harpsichord, fuzz bass, mellotron |
171
+ | `female_dark_waltz.txt` | dramatic female with vibrato over a dark 6/8 waltz, accordion, dissonant strings, tape loops (demo) |
172
+ | `female_torch_minimal.txt` | smoky low-register female torch song, felt piano, glass harmonica, long silences; use the wild recipe (rendered, not in the demos) |
173
+ | `duet_spoken_letter.txt` | the duet demo: female lead plus a male baritone narrator reading a letter |
174
+
175
+ "Female vocal" in the prompt reliably gives a female lead even though only three of the fifteen training songs have one. Waltz words ("slow waltz", "dark waltz") put the planner in 6/8; "yé-yé", "harpsichord", "fuzz bass" push it to 125 BPM in a major key.
176
+
177
+ ### A read letter (spoken passages)
178
+
179
+ YuE2 has no spoken-word mode: the planner writes a pitched melody for every vocal bar. But two things together get a passage *delivered* rather than sung:
180
+
181
+ 1. Write the passage as **prose**: long lines, no rhyme, no repetition, second person, like a letter. Put it in its own section.
182
+ 2. In the style prompt, name the narrator and say what he does: `a deep low male baritone narrator who does not sing but reads a letter aloud in a calm monotone spoken voice over the bridge, recited parlé passage with no melody ... the spoken passage close-miked and dry`.
183
+
184
+ The section tag then picks the flavour. Under a normal **`[Bridge]`** tag the planner writes the letter as a declaimed verse on a simple melody and hands it to the male voice: that is the duet demo. Under an unfamiliar **`[Spoken]`** tag the planner writes that stretch as an *instrumental interlude with no vocal melody at all*, and the decoder still voices every word over the band, at about 2.5 words per second: proper recitation. The `[Spoken]` render carried all of the letter; the `[Bridge]` render sang about half of it before returning to the chorus. Check the `.abc` next to the render: a `% interlude` with rest-only `V: Vocal` bars at that spot means it worked.
185
+
186
+ ### Lyrics
187
+
188
+ Tagged French blocks. Use `[Intro]`, `[Verse 1]`, `[Verse 2]`, `[Chorus]`, `[Bridge]`, `[Outro]` (and `[Spoken]` for narration); number the verses and write a repeated chorus out again where you want it sung. Standard orthography with accents and apostrophes. Shape that made the demos:
189
+
190
+ ```
191
+ [Intro]
192
+ two short lines, the image the song returns to
193
+
194
+ [Verse 1]
195
+ four lines, ten to twelve syllables, rhymed in pairs
196
+
197
+ [Verse 2]
198
+ four more lines
199
+
200
+ [Chorus]
201
+ four lines, the hook phrase at the top of the first two
202
+
203
+ [Verse 3]
204
+ four lines
205
+
206
+ [Chorus]
207
+ (repeat, written out)
208
+
209
+ [Bridge]
210
+ four lines, or a prose letter for the narrator
211
+
212
+ [Chorus]
213
+ (repeat, written out)
214
+
215
+ [Outro]
216
+ the intro image again, trailing off
217
+ ```
218
+
219
+ The planner sometimes drops the bridge on the orchestral prompt and goes verse–chorus–verse–chorus–outro; the acoustic and swing prompts keep it. Chanson-length lyrics like this plan to 3–5 minutes.
220
+
221
+ ### Which knob does what
222
+
223
+ - **Checkpoint** is the first choice: Rive Gauche (step 200) for anything acoustic, female, experimental or spoken; Grand Boulevard (step 350) for the big orchestral ballad.
224
+ - **strength_model (decoder)** 1.0 → 1.5 makes the voice rougher and more present without touching the writing. **strength_clip (planner) above 1.0 collapses the vocal** on YuE2 artist LoRAs in general; keep it at 1.0.
225
+ - **Weirdness (cfg)** changes only the sound. **music_temperature / score_temperature** change the writing; 1.2 / 0.9 is the wild recipe used for the torch-song prompt.
226
+ - **Seeds** decide length as much as the cap does; the same lyric planned 3:26 on one checkpoint and 4:28 on the other.
227
+
228
+ ## Training
229
+
230
+ Trained on a single RTX 5090 with the FS_Audio Suite **Artist Trainer** on Comfy's own YuE2 weights. One run trains the planner LoRA and the decoder LoRA together; the file also carries full `vae2llm` / `llm2vae` projection diffs. Both published files come from the same run.
231
+
232
+ | | Rive Gauche | Grand Boulevard |
233
+ |---|---|---|
234
+ | dataset | 15 songs, 45 minutes: French chanson from the 1950s to the 1980s, mostly orchestral male-baritone chanson with a few acoustic guitar narratives and three female-led songs | same |
235
+ | checkpoint | **step 200** | step 350 (final) |
236
+ | planner / decoder steps | 350 / 600 | 350 / 600 |
237
+ | artist (planner) loss | 5.220 → 4.606 | 5.220 → 4.582 |
238
+ | decoder loss at checkpoint | **1.148** (the minimum) | 1.162 |
239
+ | wall time | 21 minutes for the whole run, dataset build included | |
240
+
241
+ Hyper-parameters: rank 64 planner / rank 32 decoder, LR 3e-5 planner / 4e-5 decoder / 2e-5 I/O projections, artist fraction 0.5 vs regularizer pack, KL 0.1, batch 2 songs, 8192 max tokens, 750-frame windows, EMA 0.99, one song held out, score-first 0, no transcription.
242
+
243
+ **Why two checkpoints.** The trainer names its `_best` file by planner loss, which keeps falling to the end. The decoder loss bottomed at step 200 (1.148) and drifted back up to 1.162 by step 350 on this 15-song set, a mild version of the overfit seen on other small sets. Step 200 is the pick by the decoder curve, and three of the four demos are from it. The step-350 file is published because the orchestral ballad rendered from it was preferred by ear, and the drift is small enough that both are usable.
244
+
245
+ ### New-trainer run (Montmartre and Cabaret)
246
+
247
+ Same 15 songs, Ostris AI Toolkit's YuE2 trainer on the int8 ConvRot checkpoint: one rank-32 LoRA over both experts, 650 steps planned, LR 5e-5 with the planner at 0.6x, planner KL anchor 0.2, full score conditioning with 0.5 score dropout, 60 s decoder windows, planner context 4000 tokens, a checkpoint every 50 steps. About 4 s per step on an RTX 5090, 47 minutes for the run.
248
+
249
+ **Why steps 300 and 350 of 650.** With 15 songs every song is seen about 20 times by step 300. The planner's drift from the base model (the KL term) reached by step 400 what a 34-song set only reached at its final step, and by ear the decoder follows the same curve: steps 300 to 350 are clean, and from roughly step 400 on a gritty, burnt high end creeps into breaths, whispers, sibilants and cymbal swells. The later checkpoints are not published. A small set gives a usable file early, not a better one late.
250
+
251
+ ### Captioning your own dataset
252
+
253
+ The planner learns from the captions as much as from the audio, so the caption format decides whether your prompts work later. What worked here:
254
+
255
+ - **One descriptive sentence per song, trigger first**, in the order you will prompt in: language → genre → vocal → instruments → mood → production → BPM. Decade words ("1960s French chanson") are useful descriptors.
256
+ - **No tag lists, no section scaffolding** in the caption.
257
+ - **Measure BPM, don't guess**, and sanity-check it: beat trackers double slow ballads and halve fast ones (a lively yé-yé track came out as 72 and is 144; a piano ballad came out as 136 and is 68). Compare the number with the mood words in the same caption.
258
+ - A vision-language model can draft the captions from the audio, but **check the vocal gender by hand**: a high male tenor got labelled "female lead" here. The three real female leads were labelled correctly, so this is a grep-and-listen pass, not a blind find-and-replace.
259
+ - **Lyrics as tagged blocks** with standard orthography, accents and apostrophes. Transcribe at high confidence; a mis-heard lyric teaches the planner the wrong syllable count for the bar.
260
+ - **Audio prep**: FLAC 44.1 kHz / 16-bit, ≤ 320 s. MP3 rips with a damaged leading frame crash the dataset builder ("Header missing"); transcode with `ffmpeg -err_detect ignore_err`.
261
+ - **Trainer flags**: `score-first 0` and no automatic transcription.
262
+
263
+ ## Known limitations
264
+
265
+ - French vocals only.
266
+ - **High-end grit on over-trained checkpoints.** Heard on the new-trainer run from about step 400: breaths, whispered lines, "s" sounds and cymbal swells turn gritty. Montmartre (step 300) stays under it and Cabaret (step 350) shows at most a trace; if you hear it anyway, lower `strength_model` to 0.8. Part of the training audio came from lossy sources, which is a suspect we have not proven; if you train your own, start from lossless files.
267
+ - Female leads come from three training songs; they render well on the prompts above but the range of female timbres is narrower than the male ones.
268
+ - The big-band swing prompt leaves most bars instrumental. Add "sung throughout" and a denser lyric, or change the seed.
269
+ - The planner may drop a `[Bridge]` on orchestral prompts. If the bridge matters, try the acoustic prompt family or the `[Spoken]` tag.
270
+ - No instrumental-only mode is baked in; the LoRAs assume a lyric.
271
+
272
+ ## Files
273
+
274
+ ```
275
+ chnsn_montmartre.safetensors 118 MB new-trainer generation, step 300, planner + decoder LoRA (bf16, no projection diffs)
276
+ chnsn_cabaret.safetensors 118 MB new-trainer generation, step 350, same layout
277
+ chnsn_rive_gauche.safetensors 177 MB step 200, planner + decoder LoRA (bf16 weights, fp32 projection diffs)
278
+ chnsn_grand_boulevard.safetensors 177 MB step 350, same layout
279
+ demos/ mp3 renders (192 kbps from the FLAC masters), seed 7 throughout
280
+ prompts/ style prompts
281
+ ```
282
+
283
+ ## Support
284
+
285
+ These LoRAs are trained on my own GPU and released free. If they're useful to you and you'd like to chip in for compute, there's a Ko-fi: **[ko-fi.com/becausereasons](https://ko-fi.com/becausereasons)** <3
286
+
287
+ ## License and credits
288
+
289
+ Weights are released under **CC BY-NC 4.0**, inherited from the YuE2-3B base model. Non-commercial use only; attribute "CHNSN LoRAs by becausereasons".
290
+
291
+ - [YuE2](https://huggingface.co/m-a-p/YuE2-3B) by the Multimodal Art Projection (m-a-p) team; ComfyUI repack by [Comfy-Org](https://huggingface.co/Comfy-Org/YuE2).
292
+ - [ComfyUI-FS_Audio_Suite](https://github.com/KytraScript/ComfyUI-FS_Audio_Suite) by KytraScript / The Fixed Seed Company: inference nodes and the artist trainer.
293
+ - Sister releases: [CNZN canzone italiana](https://huggingface.co/becausereasons/yue2-cnzn-canzone-italiana), [MLTNT militant roots reggae](https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae).
294
+ - Trained and documented by becausereasons, September 2026.
chnsn_cabaret.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0230ff0ba9f75fcf1586b36a8dcfe67f526703c193d9a63cf1c3f5de9c7f896e
3
+ size 117500808
chnsn_montmartre.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bcbf1b5aebe2947b1677f87265f0d422adf868745bec12462e39ed9897c9dc65
3
+ size 117500808
demos/chnsn_cabaret__swing_bigband.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:efdedb7d12eff37119090e1ec6a67a672f4480eff6e68c012672bf76f935b95f
3
+ size 4941548
demos/chnsn_cabaret__yeye_baroque_clip05.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d1d35de2328e7dcb8ae3f9fba6ab61e05b5293368c298253a5581f5b80fdce3d
3
+ size 6054956
demos/chnsn_montmartre__female_dark_waltz_clip05.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:47595ab7b66248749a3b0d8bd0fcffe7f89d711f449fc4c810f213dd09afb12b
3
+ size 4914476
demos/chnsn_montmartre__female_torch.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8eea42dd299af9952100b267ea2863bc1a63b1c6db48be516849e16499da327e
3
+ size 4128236
demos/chnsn_montmartre__male_dark_crescendo.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7e1be8343111937d267f94a1bb823f7fe306fcb27c88c9b7f46756260cd82476
3
+ size 5158700
demos/chnsn_montmartre__male_piano_live.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:47a6ab7877cdad9a1541b7414a17bcc075f53e62a79f40ebb7f5d02eb516be99
3
+ size 5137388
prompts/male_dark_crescendo.txt ADDED
@@ -0,0 +1 @@
 
 
1
+ chnsn, French, dramatic French chanson, dark brooding male baritone vocal with wide rich vibrato, hushed verses rising to a full-voiced climax, acoustic piano, sweeping orchestral strings, bold brass swells, timpani, tormented and deeply melancholic, cinematic crescendo arrangement, 84 BPM
prompts/male_piano_live.txt ADDED
@@ -0,0 +1 @@
 
 
1
+ chnsn, French, French chanson, piano ballad, expressive resonant male baritone vocal with heavy trembling vibrato, hushed intimate verses and a raw belted chorus held on long sustained notes, emotive grand piano, distant orchestral strings entering late, live concert atmosphere with subtle crowd ambiance, passionate, wounded and dramatic delivery, extreme soft-to-loud dynamics, 68 BPM