SSA
Collection
Models and Datasets of paper: [Simplified Sparse Attention via Gist Tokens] • 30 items • Updated
This model is fine-tuned from gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk16 using Gist Sparse Attention (GSA) with chunk size chunk16.
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
| Field | Value |
|---|---|
| Base model | gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk16 |
| Training type | Supervised Fine-Tuning |
| Chunk size | chunk16 |
| Architecture | Llama-3.2-1B |
Base model
meta-llama/Llama-3.2-1B