--- language: - en license: apache-2.0 library_name: transformers pipeline_tag: text-generation model_name: talkie-1930-13b-it-tf base_model: - talkie-lm/talkie-1930-13b-it tags: - transformers - safetensors - bfloat16 - custom_code - text-generation - conversion - talkie - pre-1931 --- # talkie-1930-13b-it-tf (Transformers + safetensors conversion) This repository is a Transformers-compatible conversion of [`talkie-lm/talkie-1930-13b-it`](https://huggingface.co/talkie-lm/talkie-1930-13b-it), the original Talkie instruction-tuned chat model. The upstream model is an instruction-tuned post-train of `talkie-lm/talkie-1930-13b-base`, fine-tuned from instruction-response pairs extracted from pre-1931 reference works and then refined with online DPO, according to the original model card. The upstream instruction-tuned checkpoint is already BF16. This repository adds Transformers `AutoModelForCausalLM` / `AutoTokenizer` support, a chat template matching the Talkie reference code, and BF16 sharded safetensors. This is not an official Talkie release; refer to the upstream model card for the author-provided provenance and usage notes. ## Source Model - Original model: [talkie-lm/talkie-1930-13b-it](https://huggingface.co/talkie-lm/talkie-1930-13b-it) - Talkie report: [talkie-lm.com](https://talkie-lm.com/) - Reference code: [github.com/talkie-lm/talkie](https://github.com/talkie-lm/talkie) ## Conversion Details - Weight dtype: BF16 - Weight format: sharded safetensors - Context length: 4096 tokens - Architecture: custom Talkie code loaded with `trust_remote_code=True` - Tokenizer: Talkie tiktoken-compatible tokenizer exposed through `AutoTokenizer` The public reference configuration originally advertised 2,048 positions, but the Talkie team later clarified that the model was trained with a 4,096-token context. This conversion uses the corrected 4,096-token limit. ## Usage ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer path = "xlr8harder/talkie-1930-13b-it-tf" tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( path, trust_remote_code=True, dtype=torch.bfloat16, device_map={"": "cuda"}, use_safetensors=True, ) ``` For chat-style prompts: ```python messages = [{"role": "user", "content": "Write an essay predicting life in 1960."}] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_tensors="pt", return_dict=True, ).to("cuda") output = model.generate(**inputs, max_new_tokens=128) reply = output[0, inputs["input_ids"].shape[-1]:] print(tokenizer.decode(reply, skip_special_tokens=True)) ``` ## vLLM The included remote-code model implements the Transformers attention-interface hooks expected by vLLM's Transformers modeling backend. For compatibility with that backend, the original single-scalar `lm_head_gain` is folded into `lm_head.weight` during conversion; the other Talkie gain parameters remain explicit model parameters. Using vLLM's `logit_scale`-style approach was not used because it applies scaling after the output matmul, while Talkie applies the gain to the head weight before the matmul. In BF16 this can introduce small rounding differences and, in smoke tests, changed one near-tied top-token ordering. ```bash vllm serve xlr8harder/talkie-1930-13b-it-tf \ --task generate \ --model-impl transformers \ --trust-remote-code \ --dtype bfloat16 \ --max-model-len 4096 ``` ## Validation The Transformers safetensors model was compared against the original Talkie IT checkpoint on a forward-pass smoke test. The top-10 next-token ordering matched exactly; observed max absolute logit difference was `0.25`.