talkie-1930 β modern knowledge: V1 adapter ladder (talkie-timetravel)
LoRA adapter ladder from continued-pretraining talkie-lm/talkie-1930-13b-base
(13B, pre-1931 English) on real modern web text only: nvidia/ClimbMix
(docs β€1024 GPT-2 tokens, detokenized β talkie BPE) mixed 1:1 with pre-1930
replay (PG-19), ~10B tokens total across 3 fresh-data epochs. LoRA r=128 on all
7 projections (never embed/lm_head) β adapter-off is bit-identical to the
base model. Method follows the SDF paper (arXiv:2510.17941) minus synthetic
docs: this run is the real-text baseline its synthetic V2 counterpart is
compared against.
Contents
checkpoints/step_*.ptβ log-spaced adapter snapshots (fp32; the RSA/geometry ladder),checkpoints/final.ptβ endpoint (~step 76300, ~10B tokens). Load:apply_lora(model, r=128, alpha=256); load_adapter(model, ckpt["adapter"])(seetraining_code/src/talkie_cpt/lora.py).eval/offline_battery*.logβ P(fact) true/false margin battery per snapshot (450 pairs from cds-jb/talkie-timetravel-synthfactpairs), incl. obscurity tiers;eval/skyline*.jsonβ talkie-web-13b-base on the same battery.eval/geometry_ladder.jsonβ per-layer cosine + linear CKA of adapted vs base activations on held-out pre-1930 text, per snapshot.figures/β training curves (log-log losses, P(fact) error incl. obscurity tiers vs the talkie-web skyline), geometry-over-training.training_code/β the complete talkie-cpt trainer + data pipeline used to produce this run. Retrain:scripts/run_v1.sh(see its header).
Headline results (offline battery; stock β final; talkie-web skyline)
| group | stock | final | web skyline |
|---|---|---|---|
| cluster_berlin_wall | 0.378 | 0.711 | 0.889 |
| cluster_chernobyl | 0.356 | 0.644 | 0.778 |
| cluster_moon_landing | 0.422 | 0.800 | 0.889 |
| cluster_penicillin | 0.467 | 0.733 | 0.844 |
| post1930 | 0.494 | 0.806 | 0.983 |
| pre1930 | 0.922 | 0.956 | 0.978 |
Fluency: val NLL on modern text converged to the pre-1930 replay level (~1.90) after ~7B tokens. Retention: pre-1930 fact accuracy and replay NLL unchanged throughout. Knowledge by obscurity: famous facts arrive early, moderate slowly, obscure essentially not at all at this budget β the motivation for the synthetic-document V2.
License
Adapters were trained on nvidia/ClimbMix (CC-BY-NC-4.0) β this repo is CC-BY-NC-4.0. The talkie base model (Apache-2.0) is not redistributed here.
Model tree for cds-jb/talkie-1930-13b-timetravel-v1
Base model
talkie-lm/talkie-1930-13b-base