Thanks to the Talkie-1930 Team for creating such a interesting model, as well as Qwen for their contribution to the field of very small-scale language models such as Qwen3.5-0.8B.

Qwen3.5-0.8 Talkie Distilled Finetune

This model is a finetuned variation of Qwen3.5-0.8B trained on 1M tokens generated by Talkie-1930 for 3 epochs, with a rank of 8 annd Rank Alpha of 16. It was trained over thirty minutes on a RTX 5070.

Screenshot

This repository are the safetensors and configuration used to create the GGUF version of this distillation

Downloads last month
57
Safetensors
Model size
0.9B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Raydev/Qwen3.5-0.8B-Talkie-Distill-safetensors

Finetuned
(154)
this model

Dataset used to train Raydev/Qwen3.5-0.8B-Talkie-Distill-safetensors