potsawee's picture
Card: pass trust_remote_code to tokenizer load
40d4c92 verified
|
Raw
History Blame Contribute Delete
2.13 kB
metadata
license: apache-2.0
language:
  - en
tags:
  - audio
  - speech
  - next-token-prediction
  - research
  - hierarchical

p3-hier-dd512L6-5257372c

One of 13 runs from the SODA-extension flat-vs-hier campaign: a compute-matched comparison of the flattened SODA architecture (paper) against a CSM/Moshi-style hierarchical factorization (backbone over steps + small depth transformer over the 8 Mimi RVQ codebook slots), trained on the same interleaved text+audio corpus with a shared id-space so held-out NLL is directly comparable across arms.

Model details

Architecture hierarchical (backbone over steps + depth transformer over codebook slots)
Loss recipe Moshi-style alpha = 100/100/1 (text/semantic/acoustic)
Compute budget (3× forward FLOPs) 3e18
Backbone d=768, L=8
Depth transformer d=512, L=6
Window 1024 steps
Total parameters (incl. embeddings) 328M
Training step 76737
Audio Mimi RVQ, 1 semantic + 7 acoustic codebooks, 12.5 Hz

Trained on the 42k-hour campaign corpus (YODAS + Emilia EN, interleaved text+audio, single pass).

Loading

The hierarchical model ships its own modeling code (modeling_soda_hier.py, configuration_soda_hier.py) and loads with trust_remote_code:

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("soda-research/p3-hier-dd512L6-5257372c", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("soda-research/p3-hier-dd512L6-5257372c", trust_remote_code=True)