Quantized version of PatchyTisa/Qwen3.6-35B-A3B-Caption-Eagle3DraftModel.

Their model card is pasted as is below.

Files:

  • model.safetensors - original unquantized safetensors
  • model.gguf - unquantized bf16 GGUF
  • Qwen3.6-35B-A3B-Caption-Eagle3DraftModel.gguf - GGUF quantized to Q5_K_M

Eagle3 Draft Model - Qwen3.6-35B-A3B Fine-tuned Version

Model Description

This is an Eagle3 (Enhanced Auxiliary Loss for Efficient Speculative Decoding) draft model derived from the open-source Qwen/Qwen3.6-35B-A3B base model.

Training Details

  • Base Model: Qwen/Qwen3.6-35B-A3B
  • Training Dataset: Microsoft COCO train2017 caption (100k samples)
  • Training Method: Speculative decoding draft model training with Eagle3 architecture
  • Average Acceptance Rate: 83%

Model Performance

The model achieves high speculative decoding efficiency with an average acceptance rate of 83%, significantly accelerating inference when used as a draft model in speculative decoding pipelines.

Usage

This model is designed to be used as a draft model for speculative decoding with the Qwen3.6-35B-A3B target model. It can be integrated with vLLM for efficient inference acceleration.

Files

  • config.json - Model configuration file
  • model.safetensors - Model weights

Citation

If you use this model, please consider citing the Eagle3 paper and Qwen3.6 model.

Downloads last month
460
Safetensors
Model size
0.7B params
Tensor type
I64
·
BF16
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EntityDeletr/Qwen3.6-35B-A3B-Caption-Eagle3DraftModel-GGUF

Quantized
(1)
this model

Collection including EntityDeletr/Qwen3.6-35B-A3B-Caption-Eagle3DraftModel-GGUF