stream-duplex-qwen3.5-9b-stage2-245k
Stage 2 checkpoint of a streaming full-duplex speech model, at training step 245000.
This is a research checkpoint, not a drop-in transformers model. The architecture is custom:
a Qwen3.5-9B backbone with an audio branch of 8 transformer layers
forked at layer 24, 8 RVQ codebook heads
(codebook size 2048), and a frozen Kyutai Mimi encoder for user audio.
There is no AutoModel class for it, and these files cannot be loaded by name.
What is in this repo
model-*.safetensors+model.safetensors.index.json: the trainable parameters (10.85B parameters, 639 tensors). The frozen Mimi encoder and the Qwen modules the model does not use are not included; they come from the base models.meta.json: training step, stage, and the model contract the weights were trained against.
Loading
Requires the training code (StreamDuplexAudioModel), the Qwen/Qwen3.5-9B and kyutai/mimi
base weights, and a model built against the same contract as meta.json:
from safetensors.torch import load_file
import json, glob
state = {}
for shard in sorted(glob.glob("model-*.safetensors")):
state.update(load_file(shard))
model = StreamDuplexAudioModel(...) # same contract as meta.json["model_contract"]
missing, unexpected = model.load_state_dict(state, strict=False)
assert not unexpected
# `missing` should contain only the frozen Mimi encoder and unused Qwen modules.
Training
- Stage 2: backbone and audio branch trained jointly, initialized from the Stage 1 checkpoint.
- Step 245000, roughly 1.8 epochs over the training mixture.
- In-loop validation loss at this step: 1.8863, the lowest of the run so far
(lower is better; text + audio loss, see the training code for the exact definition).
meta.jsonrecordsbest_eval_loss1.8910 because a checkpoint is written before its own evaluation runs, so that field is the best loss before this step.
Data and licensing
The training mixture includes public speech datasets released under a variety of licenses, some of them restricted to non-commercial use. Check the terms of the underlying datasets before any use beyond research.