Instructions to use miguelcsx/factorized-natural-dense with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use miguelcsx/factorized-natural-dense with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="miguelcsx/factorized-natural-dense", trust_remote_code=True)# Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("miguelcsx/factorized-natural-dense", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
FACTORIZED Natural Dense
Selected strict-small checkpoint from the controlled FACTORIZED study. The model uses an LTG-BERT-style encoder trained with whole-word masking and an auxiliary replaced-token objective. Its 384-dimensional representation is partitioned into three 128-dimensional factor readouts. In this natural-dense control the factor readout is dormant during language-model inference.
Model summary
| Property | Value |
|---|---|
| Parameters | approximately 33.3M |
| Layers / hidden size | 12 / 384 |
| Attention heads | 6 |
| FFN size | 1,280 |
| Context length | 512 |
| Tokenizer | byte-level BPE, 16,384 tokens |
| Training corpus | BabyLM strict-small, conservatively counted as 10M words |
| Total exposure | 100M words |
| Optimizer | LAMB |
| Peak learning rate | 3.5e-3 |
Usage
from transformers import AutoModelForMaskedLM, AutoTokenizer
repo_id = "miguelcsx/factorized-natural-dense"
revision = "chck_100M"
tokenizer = AutoTokenizer.from_pretrained(repo_id, revision=revision)
model = AutoModelForMaskedLM.from_pretrained(
repo_id,
revision=revision,
trust_remote_code=True,
)
Remote code is required for the custom TOLM architecture. Review tolm.py
before loading it in a security-sensitive environment.
Development evaluation
The selected chck_100M checkpoint led the FACTORIZED tournament on the
pre-registered shared full-task mean among the three finalists.
| Evaluation | Score |
|---|---|
| BLiMP | 70.01 |
| BLiMP Supplement | 60.63 |
| Entity Tracking | 17.70 |
| COMPS | 51.41 |
| GlobalPIQA | 36.09 |
These are development measurements, not a claim of an official leaderboard placement. The complete official BabyLM 2026 evaluation is being regenerated for submission.
Checkpoint revisions
The repository contains all strict-small challenge revisions:
chck_1M through chck_10M, then chck_20M through chck_100M.
The final submission checkpoint is chck_100M.
Provenance
- Final checkpoint model SHA-256:
ba89940028a902f4aa09385966b8e83bb346a0abaab6899feec1fa2d3a4b7e13. - Evaluation backend:
mntp. - Fast-evaluation parity maximum absolute logit delta:
0.0. - Training and compliance details are recorded in
training_manifest.json.
Limitations
This is a small English developmental language model trained for controlled BabyLM experiments. It is not a general-purpose assistant. The factorized readout in this control is dormant during ordinary masked-language-model inference.
- Downloads last month
- 186