--- library_name: transformers pipeline_tag: text-generation license: apache-2.0 language: - "en" - "zh" - "fr" - "es" - "de" - "pt" - "ru" - "it" - "ja" - "ko" - "vi" - "ar" datasets: - "HuggingFaceFW/fineweb-edu" - "mlfoundations/dclm-baseline-1.0" - "cerebras/SlimPajama-627B" - "EleutherAI/pile" - "bigcode/starcoderdata" - "oscar-corpus/OSCAR-2301" tags: - rwkv - rwkv7 - recurrent - causal-lm - conversational ---
RWKV logo

RWKV7-G1j-13.3B-20260831

RWKV-7 “Goose” · constant-state recurrent language modeling

Website Hugging Face GitHub RWKV-7 paper License
--- ## Model introduction This is an official BlinkDL release of **RWKV-7 Goose** in Hugging Face Transformers format. RWKV-7 is an attention-free recurrent architecture with a constant-size recurrent state and constant inference work per generated token. Training remains parallelizable. This checkpoint is a **base model** pretrained with web, code, synthetic, instruction, chat, and reasoning data. It is suitable for evaluation, post-training, and fine-tuning; the included chat template is a prompt interface, not a claim that the checkpoint is a safety-aligned assistant. The Transformers integration, conversion, release packaging, linear-time RWKV tokenizer, and optional TileLang inference implementation are distributed with this release. ## Highlights - **Constant recurrent state:** memory does not grow like an attention KV cache. - **Bundled Transformers integration:** auditable remote configuration and modeling modules provide generation, recurrent cache continuation, training, and LoRA workflows on Transformers 5.15+. - **Exact linear-time tokenizer:** the bundled RWKV trie reads the self-contained `tokenizer.json` generated from the canonical RWKV World byte vocabulary. - **Chat-ready:** `chat_template.jinja` supports system, multi-turn, thinking, and strict model-generated tool-call prompts. - **Optional optimized runtime:** the isolated [`inference/`](inference/) bundle provides PyTorch fallback and TileLang acceleration without changing the standard model root. ## Model overview | Field | Value | | --- | --- | | Repository | `RWKV/RWKV7-G1j-13.3B-20260831` | | Architecture class | `Rwkv7ForCausalLM` | | Public size label | `13.3B` | | Source parameters | `13,270,298,624` | | Serialized parameters | `13,270,298,624` | | Synthesized compatibility tensors | `0` | | Layers | `61` | | Hidden / FFN size | `4096` / `16384` | | Heads / head size | `64` / `64` | | Vocabulary | `65536` | | Training context | `16384 tokens` | | Weight dtype | `bfloat16` | | Numerical conversion | `source dtype preserved` | | Metadata profile | `g1j` | | Metadata provenance | `locked-profile` | | Source checkpoint | [`BlinkDL/rwkv7-g1/rwkv7-g1j-13.3b-20260831-ctx16384.pth`](https://huggingface.co/BlinkDL/rwkv7-g1/blob/fd65209af38473e856ae69422c895814bbcdd2a6/rwkv7-g1j-13.3b-20260831-ctx16384.pth) | | Source SHA-256 | `559371f5b9aef13189ae54b345ac096af4ad2b689996c05d89de687612b3ae65` | ## Transformers quickstart The repository includes `configuration_rwkv7.py`, `modeling_rwkv7.py`, and the exact linear-time `tokenization_rwkv7.py`. The model modules are adapted from the Transformers RWKV-7 integration at commit [`4ad9ed0`](https://github.com/huggingface/transformers/commit/4ad9ed0747ed6ba75c787e8f9040dcd64b166ee2). Review those files and pin a model-repository revision in production. Passing `trust_remote_code=True` selects this bundled implementation even when the local Transformers installation also provides native RWKV-7 support. ```python import torch from transformers import ( AutoModelForCausalLM, AutoTokenizer, ) model_id = "RWKV/RWKV7-G1j-13.3B-20260831" tokenizer = AutoTokenizer.from_pretrained( model_id, trust_remote_code=True, ) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, dtype=torch.bfloat16, ) ``` The recurrent cache returned by the model can be passed back for incremental decoding. Use an `attention_mask` for padded batches. The model defaults to the chunk-parallel WKV path for efficient multi-token prefill. To reproduce the portable token-order reference path, set `model.config.wkv_implementation = "eager"` before the first forward pass. Chunked execution changes floating-point operation order, so small numerical differences from eager execution are expected. ## Chat quickstart ```python import re import torch from transformers import AutoModelForCausalLM, AutoTokenizer THINK_RE = re.compile(r"\A?\s*(.*?)\s*?", re.DOTALL) def assistant_content(completion, thinking, *, close_incomplete=False): prefix = "\n" reply = prefix + completion thinking_block = THINK_RE.match(reply) if thinking: if thinking_block is not None or not close_incomplete: return reply.strip() return f"{reply.rstrip()}\n".strip() return "" if thinking_block is None else reply[thinking_block.end():].strip() model_id = "RWKV/RWKV7-G1j-13.3B-20260831" tokenizer = AutoTokenizer.from_pretrained( model_id, trust_remote_code=True, ) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, dtype=torch.bfloat16, ).to("cuda") messages = [{"role": "user", "content": "Explain why RWKV uses constant state."}] thinking = False max_new_tokens = 256 inputs = tokenizer.apply_chat_template( messages, tokenize=True, add_generation_prompt=True, thinking=thinking, return_dict=True, return_tensors="pt", ).to(model.device) output = model.generate( **inputs, max_new_tokens=max_new_tokens, do_sample=True, temperature=1.0, top_p=0.5, eos_token_id=0, pad_token_id=0, stop_strings=["\n\nUser:"], tokenizer=tokenizer, ) completion = tokenizer.decode( output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True, ) completion = completion.split("\n\nUser:", 1)[0] reached_token_limit = output.shape[1] - inputs["input_ids"].shape[1] >= max_new_tokens print( assistant_content( completion, thinking, close_incomplete=reached_token_limit, ) ) ``` Set `thinking=True` for the RWKV thinking prefix. The intentional generation prefixes are `Assistant: ` followed by a newline and `Assistant: