File size: 1,253 Bytes
5befb4b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 | ---
license: apache-2.0
library_name: pytorch
tags:
- image-classification
- cifar-10
- lra
- sequence-model
- interdomain-attention
datasets:
- cifar10
metrics:
- accuracy
---
# LRA-Image Softmax (RoPE, causal)
Softmax (RoPE, causal) model trained on the Long Range Arena (LRA) sCIFAR-10 benchmark.
## Model Details
- **Architecture**: Softmax (RoPE, causal)
- **Task**: Sequential CIFAR-10 (grayscale, 1024 tokens)
- **Parameters**: ~4.1M
- **Position encoding**: RoPE
- **Causal**: True
- **Test Accuracy**: 68.88%
## Training
- **Protocol**: S4 paper LRA-Image (200 epochs, lr=1e-3, batch=64, warmup=18k steps)
- **Backbone**: Llama-style (6 layers, d=512, 8 heads, RMSNorm, SwiGLU)
- **Seed**: 2222
- **WandB run**: [harrisonzhu/InterdomainAttention/x437k0ed](https://wandb.ai/harrisonzhu/InterdomainAttention/runs/x437k0ed)
## Usage
```python
import torch
from model import LlamaLRAImage # requires interdomain-attention repo
state_dict = torch.load("lra_image_best.pt", weights_only=True)
model = LlamaLRAImage(...) # match config
model.load_state_dict(state_dict["model"])
```
## Citation
```bibtex
@article{interdomain2026,
title={Interdomain Attention},
author={...},
year={2026}
}
```
## License
Apache 2.0
|