talkie-1930-13b-base — J-Lens (shared pre-1931 bridge)

This repository contains a fitted Jacobian lens for talkie-lm/talkie-1930-13b-base, a 13B base model trained on pre-1931 English-language text. It contains lens weights only; the 13B base-model weights are downloaded separately from the linked Talkie repository.

For each source layer l, the lens stores the corpus-average Jacobian J_l that maps a residual-stream vector into the final-layer basis before applying the model's own unembedding. This is an observational interpretability instrument, not a classifier and not a claim about model beliefs, intent, truth, or emotion.

Compatibility

Field Value
Base model talkie-lm/talkie-1930-13b-base
Fitted base revision b7c97680791f7fca4262c3c80b36ff7d666faab0
Residual width 5120
Source layers 27–34
Target layer 39
Matrices 8 × 5120 × 5120
Stored dtype float16 (loaded as float32 by jlens)
Fit prompts 233 × 128 tokens
Lens SHA-256 ab81abad5a7e4b44cb05f8f28eb682ff421c8ae31a81fdfa01de66bb56b21eb9

The lens is architecture- and revision-specific. Do not apply it to another Talkie checkpoint, an instruction-tuned variant, or a model with the same width but different weights.

Load the lens

Install the pinned reference implementations:

pip install -r requirements.txt

Then load from the Hub:

import jlens

lens = jlens.JacobianLens.from_pretrained(
    "osoleve/talkie-1930-13b-base-jlens-pre1931-bridge"
)
print(lens)

For a local clone, JacobianLens.from_pretrained(".") works because the weight file uses the library's default lens.pt name.

example_usage.py can inspect the package without loading the base model:

python example_usage.py

To apply the lens at the last prompt position (requires a CUDA GPU with enough memory for Talkie 13B):

python example_usage.py --prompt "The effects of the automobile on public morality have"

The included talkie_adapter.py is the minimal LensModel adapter used for fitting and application because Talkie uses its own native model class rather than a Hugging Face Transformers class.

Fit corpus

Both published Talkie lenses were fitted independently on the same 233 raw passages from a balanced, public-domain 1880–1930 bridge corpus. The exact prompts and item-level provenance are included in fit_corpus.json.

Source Passages
Project Gutenberg 39
HathiTrust / Internet Archive books 39
Library of Congress books 39
American Stories newspapers 39
USPTO patents 39
U.S. case law 38

Prompt-list SHA-256: a270a07027e7ee6875ac2392a7a689116f9398deca5f85b3a11926bd67db8693. The bridge is common support for a controlled comparison; it is not a reconstruction of Talkie's unreleased pretraining mixture and does not estimate that mixture's source weights.

Fit configuration

Field Value
Jacobian dimension batch 64
Maximum sequence length 128
GPU NVIDIA RTX PRO 6000 Blackwell Server Edition
PyTorch 2.11.0+cu128
Elapsed fit time 94.0 minutes
Talkie code revision 35317ba3a84861a84c84065bd73faf88ad19329c
Jacobian Lens revision 581d398613e5602a5af361e1c34d3a92ea82ba8e

Validation

Every stored matrix was checked for shape 5120 × 5120 and finite values. The paired bridge-corpus lenses have mean matrix cosine similarity 0.649826 and mean relative Frobenius divergence 0.838331 across layers 27–34. Full per-layer results are in validation.json; the bridge-vs-WikiText 2×2 control is in comparison_2x2.json.

Run the self-contained integrity check after cloning:

python verify.py

Files

File Purpose
lens.pt jlens.JacobianLens checkpoint
lens_config.json Portable compatibility and provenance manifest
fit_metadata.json Original fit settings, with portable paths
fit_corpus.json Exact prompts, corpus statistics, and source revisions
validation.json Paired matrix integrity and divergence results
comparison_2x2.json Bridge-corpus vs. WikiText control comparison
talkie_adapter.py Native Talkie LensModel adapter
example_usage.py Local/Hub loading and application example
verify.py Hash, corpus, shape, and finiteness verification
SHA256SUMS Package file checksums

Limitations and responsible use

  • Readouts depend on the fit corpus, model revision, layer, token position, and the model's unembedding; interpret them as exploratory evidence.
  • The 233-prompt fit is sufficient for a useful research lens but is not an exhaustive estimator of every domain or register.
  • Differences between the vintage and web models cannot be attributed only to time: their pretraining distributions also differ in subject matter and other ways.
  • These are unfiltered base language models and may produce explicit, offensive, stereotyped, or otherwise disturbing text. The lens does not make generation safer.

License and attribution

The fitted lens weights and package support files are released under Apache License 2.0. The Talkie base model and reference implementation are Apache-2.0, as is Anthropic's Jacobian Lens reference implementation. See LICENSE and NOTICE.md. No Talkie base-model weights are redistributed here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for osoleve/talkie-1930-13b-base-jlens-pre1931-bridge

Finetuned
(13)
this model

Datasets used to train osoleve/talkie-1930-13b-base-jlens-pre1931-bridge

Collection including osoleve/talkie-1930-13b-base-jlens-pre1931-bridge