nkiyohara's picture
Upload LRA-Image Softmax (RoPE, causal) (test_acc=68.88%)
5befb4b verified
|
Raw
History Blame Contribute Delete
1.25 kB
metadata
license: apache-2.0
library_name: pytorch
tags:
  - image-classification
  - cifar-10
  - lra
  - sequence-model
  - interdomain-attention
datasets:
  - cifar10
metrics:
  - accuracy

LRA-Image Softmax (RoPE, causal)

Softmax (RoPE, causal) model trained on the Long Range Arena (LRA) sCIFAR-10 benchmark.

Model Details

  • Architecture: Softmax (RoPE, causal)
  • Task: Sequential CIFAR-10 (grayscale, 1024 tokens)
  • Parameters: ~4.1M
  • Position encoding: RoPE
  • Causal: True
  • Test Accuracy: 68.88%

Training

  • Protocol: S4 paper LRA-Image (200 epochs, lr=1e-3, batch=64, warmup=18k steps)
  • Backbone: Llama-style (6 layers, d=512, 8 heads, RMSNorm, SwiGLU)
  • Seed: 2222
  • WandB run: harrisonzhu/InterdomainAttention/x437k0ed

Usage

import torch
from model import LlamaLRAImage  # requires interdomain-attention repo

state_dict = torch.load("lra_image_best.pt", weights_only=True)
model = LlamaLRAImage(...)  # match config
model.load_state_dict(state_dict["model"])

Citation

@article{interdomain2026,
  title={Interdomain Attention},
  author={...},
  year={2026}
}

License

Apache 2.0