davidtai-eigenlabs's picture
Model card: simplified technical English
e678963 verified
|
Raw
History Blame Contribute Delete
1.45 kB
---
language: en
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
tags:
- mlx
- qwen3_5_mtp
- speculative-decoding
---
# Qwen3.8-27B-MTP-bf16
This is the MTP (multi-token-prediction) head of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B).
The official model contains this head as 15 `mtp.*` tensors.
We copied these 15 tensors from revision `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0`.
We removed the `mtp.` prefix from each tensor name. Example: `mtp.fc.weight` is now `fc.weight`.
We did not train the weights. We did not change the weights.
The data type is bfloat16, the same as the official release.
## Contents
- One fusion projection (`fc.weight`).
- One transformer layer with its norms.
- `config.json` with `model_type: qwen3_5_mtp` and `block_size: 3`.
- `model.safetensors.index.json`. Loaders read the head as its own model tree, so these two files are necessary.
## This is not a standalone model
The head predicts draft tokens from the hidden states of its base model.
It cannot generate text alone.
Use it with the 4-bit reference model:
[EigenLabs/Qwen3.8-27B-4bit](https://huggingface.co/EigenLabs/Qwen3.8-27B-4bit).
## Role
The Qwen 3.8 MTP challenge uses this head for speculative decoding.
The track is `qwen3.8-27b-mtp-v1`.
The challenge repository is [Layr-Labs/qwen-3.8-mtp-challenge](https://github.com/Layr-Labs/qwen-3.8-mtp-challenge).
The challenge points to one fixed revision of this repository.