File size: 1,253 Bytes
5befb4b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
---
license: apache-2.0
library_name: pytorch
tags:
  - image-classification
  - cifar-10
  - lra
  - sequence-model
  - interdomain-attention
datasets:
  - cifar10
metrics:
  - accuracy
---

# LRA-Image Softmax (RoPE, causal)

Softmax (RoPE, causal) model trained on the Long Range Arena (LRA) sCIFAR-10 benchmark.

## Model Details

- **Architecture**: Softmax (RoPE, causal)
- **Task**: Sequential CIFAR-10 (grayscale, 1024 tokens)
- **Parameters**: ~4.1M
- **Position encoding**: RoPE
- **Causal**: True
- **Test Accuracy**: 68.88%

## Training

- **Protocol**: S4 paper LRA-Image (200 epochs, lr=1e-3, batch=64, warmup=18k steps)
- **Backbone**: Llama-style (6 layers, d=512, 8 heads, RMSNorm, SwiGLU)
- **Seed**: 2222
- **WandB run**: [harrisonzhu/InterdomainAttention/x437k0ed](https://wandb.ai/harrisonzhu/InterdomainAttention/runs/x437k0ed)

## Usage

```python
import torch
from model import LlamaLRAImage  # requires interdomain-attention repo

state_dict = torch.load("lra_image_best.pt", weights_only=True)
model = LlamaLRAImage(...)  # match config
model.load_state_dict(state_dict["model"])
```

## Citation

```bibtex
@article{interdomain2026,
  title={Interdomain Attention},
  author={...},
  year={2026}
}
```

## License

Apache 2.0