talkie-1930 β†’ modern knowledge: V1 adapter ladder (talkie-timetravel)

LoRA adapter ladder from continued-pretraining talkie-lm/talkie-1930-13b-base (13B, pre-1931 English) on real modern web text only: nvidia/ClimbMix (docs ≀1024 GPT-2 tokens, detokenized β†’ talkie BPE) mixed 1:1 with pre-1930 replay (PG-19), ~10B tokens total across 3 fresh-data epochs. LoRA r=128 on all 7 projections (never embed/lm_head) β€” adapter-off is bit-identical to the base model. Method follows the SDF paper (arXiv:2510.17941) minus synthetic docs: this run is the real-text baseline its synthetic V2 counterpart is compared against.

Contents

  • checkpoints/step_*.pt β€” log-spaced adapter snapshots (fp32; the RSA/geometry ladder), checkpoints/final.pt β€” endpoint (~step 76300, ~10B tokens). Load: apply_lora(model, r=128, alpha=256); load_adapter(model, ckpt["adapter"]) (see training_code/src/talkie_cpt/lora.py).
  • eval/offline_battery*.log β€” P(fact) true/false margin battery per snapshot (450 pairs from cds-jb/talkie-timetravel-synth factpairs), incl. obscurity tiers; eval/skyline*.json β€” talkie-web-13b-base on the same battery.
  • eval/geometry_ladder.json β€” per-layer cosine + linear CKA of adapted vs base activations on held-out pre-1930 text, per snapshot.
  • figures/ β€” training curves (log-log losses, P(fact) error incl. obscurity tiers vs the talkie-web skyline), geometry-over-training.
  • training_code/ β€” the complete talkie-cpt trainer + data pipeline used to produce this run. Retrain: scripts/run_v1.sh (see its header).

Headline results (offline battery; stock β†’ final; talkie-web skyline)

group stock final web skyline
cluster_berlin_wall 0.378 0.711 0.889
cluster_chernobyl 0.356 0.644 0.778
cluster_moon_landing 0.422 0.800 0.889
cluster_penicillin 0.467 0.733 0.844
post1930 0.494 0.806 0.983
pre1930 0.922 0.956 0.978

Fluency: val NLL on modern text converged to the pre-1930 replay level (~1.90) after ~7B tokens. Retention: pre-1930 fact accuracy and replay NLL unchanged throughout. Knowledge by obscurity: famous facts arrive early, moderate slowly, obscure essentially not at all at this budget β€” the motivation for the synthetic-document V2.

License

Adapters were trained on nvidia/ClimbMix (CC-BY-NC-4.0) β†’ this repo is CC-BY-NC-4.0. The talkie base model (Apache-2.0) is not redistributed here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for cds-jb/talkie-1930-13b-timetravel-v1

Adapter
(3)
this model

Paper for cds-jb/talkie-1930-13b-timetravel-v1