GSA-FT-Llama-3.2-1B-chunk16

This model is fine-tuned from gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk16 using Gist Sparse Attention (GSA) with chunk size chunk16.

Paper

GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding

Model Details

Field Value
Base model gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk16
Training type Supervised Fine-Tuning
Chunk size chunk16
Architecture Llama-3.2-1B
Downloads last month
17
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gist-sparse-attention/GSA-FT-Llama-3.2-1B-chunk16

Collection including gist-sparse-attention/GSA-FT-Llama-3.2-1B-chunk16