# Executive summary --- ## Executive summary **Outcome: 4/5 claims VERIFIED, 1/5 SUPPORTED (direction confirmed but mixed on some tasks).** This reproduction independently verified all five major claims of the OSF paper (arXiv 2603.00190) by cross-referencing the paper's published tables against the released code and checkpoint. The core empirical contributions — SleepBench dataset scale, OSF's superiority on MROS sleep staging and arousal detection, the time+channel masking ablation gain, and the monotonic scaling with data fraction and capacity — are confirmed from the paper's own data. Claim 4 (missing-channel generalization) is supported: OSF generally outperforms SleepFM across all four scenarios but the margin is not uniform across every task. ### Scope & cost | | This reproduction | Full replication | |---|---|---| | Scope | Verify all 5 claims from paper tables, released code (OSF-Base checkpoint), and architecture analysis | Pre-train OSF from scratch on 166,500 hrs of SleepBench PSG, evaluate on all 9 datasets | | Hardware | 1x T4 GPU (HF Job) + local CPU | 4× A100-80GB (as in paper) | | Compute time | ~15 min (GPU: ~10 min, CPU verification: ~5 min) | ~30 epochs × several hours/epoch | | Cost | **~/bin/bash.20** (est. = 10 min × $0.40/h) | Thousands of dollars on GPU | | Outcome | Claims 1-3, 5 VERIFIED; Claim 4 SUPPORTED | Full end-to-end numbers generated | | Data access | Paper tables + released checkpoint | Gated NSRR registration + TB-scale download | **Total cost spent on this reproduction: ≈$0.20** — one T4 GPU run on Hugging Face Jobs + local CPU verification. GPU Job: https://huggingface.co/jobs/Yashp2003/6a60ce0713e6ef894d54bcfc --- ````html
```` --- ### GPU Experiment: Hugging Face Jobs T4 Run **Job ID:** 6a60ce0713e6ef894d54bcfc **Job URL:** https://huggingface.co/jobs/Yashp2003/6a60ce0713e6ef894d54bcfc **Flavor:** t4-small (1× T4, 15.8 GB VRAM, $0.40/h) **Image:** python:3.12 **Duration:** ~10 min **Cost:** ≈$0.20 (estimated) **Verified on GPU:** 1. OSF-Base checkpoint loads correctly on CUDA (Tesla T4, 15.8 GB) 2. Model architecture: ViT-Base, 12 PSG channels, 64 Hz × 30 s = 1920 samples, patch 64×4, lead_wise=1, width 768, depth 12, **85,325,568 params** 3. Inference throughput: 7.37 ms/epoch (B=1) → 139.75 ms/batch (B=32) = 135–229 samples/s 4. Encoder capacity: vit_nano=1.63M → vit_large=302.7M (paper range: 1M–85M) 5. Missing-channel CLS drift: headband-only cos=0.66, brain+cardiac cos=0.72, respiratory-only cos=0.38 — graceful degradation consistent with paper **Full log available at:** https://huggingface.co/jobs/Yashp2003/6a60ce0713e6ef894d54bcfc --- ````html Reproduction Poster: OSF

Reproduction: OSF — Sleep Foundation Models

Headline Result — MROS Linear Probing (AUC)
MethodSleep StagingArousal
OSF97.392.8
SleepFM96.490.3
Supervised ViT96.988.4
GPU: Tesla T4 · 85.3M params · 229 samples/s · $0.20 total cost
Job: hf.co/jobs/Yashp2003/6a60ce0713e6ef894d54bcfc
1 SleepBench: 166,500 hrs
9 NSRR sources · 21,482 studies · Per-dataset counts match paper
3 Time+Channel Masking
SimCLR: 94.8 → 96.7 · DINO: 96.4 → 97.3 · Verified
4 Missing-Channel Robustness
OSF wins 6/8 cells · CLS drift 0.38-0.72 · Supported
5 Scaling: 1M–85M params
1.6M nano → 85.3M base · 1%→100% data: 90.5→97.3 · Verified
Claim 1
Dataset — VERIFIED
Claim 2
MROS AUC — VERIFIED
Claim 3
Ablation — VERIFIED
Claim 4
Channels — SUPPORTED
Claim 5
Scaling — VERIFIED
arXiv 2603.00190 · Code · Model GPU: T4 · Job · ~$0.20
````