Raydev/Talkie1930-1M
Viewer • Updated • 3.29k • 56
Thanks to the Talkie-1930 Team for creating such a interesting model, as well as Qwen for their contribution to the field of very small-scale language models such as Qwen3.5-0.8B.
This model is a finetuned variation of Qwen3.5-0.8B trained on 1M tokens generated by Talkie-1930 for 3 epochs, with a rank of 8 annd Rank Alpha of 16. It was trained over thirty minutes on a RTX 5070.
This repository are the safetensors and configuration used to create the GGUF version of this distillation