CSC 509 avGFP ESM-2 150M Full-Dataset Fine-Tune
This model is a full fine-tune of facebook/esm2_t30_150M_UR50D for avGFP fluorescence regression.
Training data
- Source CSV:
../data/GFP_AEQVI_Sarkisyan_2016.csv - Rows used for training: 51,714
- Target column:
DMS_score - Input column:
mutated_sequence - Source CSV SHA-256:
05b97d36a23363331b271edc56f1496b9a539109fd006e65532f91e9903a0838
Important caveat
This checkpoint was trained on the full avGFP table. If the W6 lecture split file was present,
that means the W6/W7/W8 held-out rows were included in training. Evaluation on that held-out
partition should be labeled full_INFLATED in W8L3.
Hyperparameters
- Epochs: 5
- Batch size: 4
- Gradient accumulation steps: 1
- Effective batch size: 4
- Learning rate: 2e-05
- Weight decay: 0.0
- Max token length: 512
- Mixed precision: True
- Seed: 509
W8L3 held-out Spearman
0.8711
- Downloads last month
- 6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for dsml-biotech-cert/csc509-avgfp-esm2-150m-finetune-full
Base model
facebook/esm2_t30_150M_UR50D