--- license: agpl-3.0 base_model: Qwen/Qwen3-4B-Base pipeline_tag: text-generation language: - en tags: - prose - rewriting - style-transfer - creative-writing - deslop - qwen3 library_name: transformers --- # prose-rewriter-4b-v1.4 A paragraph-level **prose rewriter**: it takes prose written by a large model and re-renders it to be more human, preserving the semantics it was given. `Qwen/Qwen3-4B-Base` with a rank-32 LoRA merged in at strength 1.20. Successor to [prose-rewriter-4b-v1.3](https://huggingface.co/chartreuse-verte/prose-rewriter-4b-v1.3). It **edits where v1.3 passed** -- it hands a paragraph back essentially unchanged a third as often, and varies sentence length about twice as much -- on a pool that adds a roleplay-forum register, and a negation slop banishment (e.g. `He didn't answer.`). See [Evaluation](#evaluation). ## Variants | Path | Format | Use with | |---|---|---| | `/` | safetensors bf16, `qwen3` arch | transformers | | `GGUF/prose-rewriter-4b-v1.4-Q8_0.gguf` | GGUF Q8_0, 4.69 GB | llama.cpp / llama-cpp-python | | `GGUF/prose-rewriter-4b-v1.4-Q4_K_M.gguf` | GGUF Q4_K_M, 2.72 GB | llama.cpp / llama-cpp-python | The quants carry the chat template and stop on `<|im_end|>`, and the adapted output head is kept separate from the token embeddings in both — Q8_0 stores it at Q8_0, Q4_K_M at Q6_K. ## Prompt format ``` <|im_start|>source {paragraph}<|im_end|> <|im_start|>edit match<|im_end|> <|im_start|>rewrite ``` The chat template in this repo builds exactly that string, byte for byte, from two roles: ```python messages = [ {"role": "source", "content": paragraph}, {"role": "edit", "content": "match"}, ] tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) ``` It is not a chat model. The template rejects an `edit` value outside the three modes rather than quietly building a prompt the weights have never seen. Any other role is treated as the source paragraph, so a runtime that probes the template with a `user` message still gets a valid prompt. ## The `edit` block is mandatory `edit` names which of three length transforms is being asked for. **The values describe the input, not the instruction.** They say what kind of text you are handing over: | `edit` | what it says about the input | what the model does | |---|---|---| | `match` | the source is about the length it should be | rewrite in place | | `inflate` | the source is padded relative to what it should be | cut | | `compress` | the source is flattened and too short | open it back out | `match` is the setting for "rewrite it, do not trim it". It's strongly recommended you use this mode. Sending **no** block is the worst thing you can do to this checkpoint. It was trained with the block, so omitting it collapses the model onto its deletion-heaviest mode. ## Serving recipe Sampled at `temperature=0.9, top_p=0.9`. Temperature 0.9 has been tested and internally to be the most optimal value. It's recommended you use this. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo = "chartreuse-verte/prose-rewriter-4b-v1.4" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="cuda").eval() def rewrite(paragraph, mode="match"): text = tok.apply_chat_template( [{"role": "source", "content": paragraph}, {"role": "edit", "content": mode}], tokenize=False, add_generation_prompt=True, ) ids = tok(text, return_tensors="pt", add_special_tokens=False).input_ids.to(model.device) out = model.generate(ids, max_new_tokens=512, do_sample=True, temperature=0.9, top_p=0.9) return tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True).strip() ``` Same thing under llama.cpp. The roles are `source` and `edit`, which no chat API models, so build the string yourself; `<|im_end|>` stops it: ```bash llama-cli -m GGUF/prose-rewriter-4b-v1.4-Q8_0.gguf -no-cnv -n 512 --temp 0.9 --top-p 0.9 \ -p '<|im_start|>source {paragraph}<|im_end|> <|im_start|>edit match<|im_end|> <|im_start|>rewrite ' ``` ## Input length The training pool's median input is 42 words and most of it is under 70, so serve it on anything from a full sentence up. The practical floor is about **15 words**. Below it the failure mode is padding and fabrication rather than gibberish: the model stretches the line toward its learned length and adds material the input never supported. Below 80 bytes, pass the text through unchanged. ## Evaluation Both releases measured **as they ship** -- v1.3 baked at strength 1.2, v1.4 at 1.20 -- on 365 held-out paragraphs of LLM-written prose that neither model saw in training, both sent the identical prompt at `temperature=0.9, top_p=0.9`, three swipes each. Paired over the 356 inputs of 51 words or more. | | v1.3 | v1.4 | paired *t* | |---|---|---|---| | words changed | 29.9% | **36.6%** | +9.86 | | passed through unchanged | 10.3% | **3.4%** | -5.66 | | below the training edit floor | 71.5% | **53.7%** | -9.73 | | sentence-length variety vs input | +0.061 | **+0.113** | +7.76 | | sentence count moved | 69.9% | **77.2%** | +3.82 | | words kept from the input | 0.734 | 0.681 | -10.67 | | truncated below 0.75x | 11.2% | 14.4% | +2.43 | | length preserved | 0.883 | 0.881 | -0.44 | | near-verbatim outputs | 2.3% | 1.5% | -1.21 | | repeated 3-grams | 0.005 | 0.004 | -0.78 | **v1.4 edits where v1.3 passed.** It hands a paragraph back essentially untouched on 3.4% of inputs against v1.3's 10.3%, and clears the edit floor its own training pairs were gated at on a third more paragraphs. The largest single move is sentence-length variety, which roughly doubles: this model breaks and joins sentences rather than swapping words inside the ones it was given, and the sentence-count column agrees. **It also truncates more, and that one is real.** Dropping below 0.75x the input's length rises 3.2 points, 95% [+0.6, +5.8]. Overall length preservation does not move (-0.44, n.s.), so this is a tail rather than a shift -- but it is the cost of the strength this release is baked at, and it is the number to watch if you feed it long paragraphs. Repetition and near-verbatim copying do not move. The register numbers on the same paragraphs, this time unpaired and against the input rather than against v1.3: | | input | v1.4 | human corpus | |---|---|---|---| | banned constructions /1k | 7.52 | **2.72** | 0.00 | | slop lexicon density | 0.090 | **0.057** | 0.021 | | purple score | 0.498 | **0.320** | -- | ### A note on strength **The strength numbers on this card and on v1.3's are not comparable.** LoRA scaling is `(alpha / r) * lora_scale`, and alpha changed between the two runs: v1.3's adapter was alpha 32 at rank 32, so its `1.2` was an effective 1.2. This one is alpha 64 at rank 32, so the same `1.2` is an effective **2.40** -- twice the multiplier on a delta that was trained under a different alpha. If you have been running v1.3 at some favourite strength, do not carry the number across; find this checkpoint's own. Every number above is measured at the strength this release is baked at rather than at the adapter's natural 1.0, because those differ and the merge is what ships. Read at 1.0 this same checkpoint changes about 30% of words and passes about 19% of paragraphs through (n=80 strength sweep), which describes the training run and not this artifact. ## Training **Corrupt forward, train backward.** The human paragraph is the target; an on-policy LLM manufactures the input by slop-ifying it. The target side is human prose: roughly half r/WritingPrompts ([`Mollymo/Human-to-AI-writing`](https://huggingface.co/datasets/Mollymo/Human-to-AI-writing)) and two-fifths AO3 ([`midwestern-simulation-active/ao3_random_subset`](https://huggingface.co/datasets/midwestern-simulation-active/ao3_random_subset)), with a sliver of fanfiction.net ([`atom-in-the-universe/fanfics-10k-10k`](https://huggingface.co/datasets/atom-in-the-universe/fanfics-10k-10k)). **New since v1.3: a roleplay-forum register.** Everything above is prose fiction, and this rewriter deploys on roleplay. A scrape of `bluemoonroleplaying.com` now supplies 9.3% of the target side — the only human writing in the pool that is already in the deployment's own register. The input side is manufactured from those targets, weighted and share-capped so that no single generator's tics dominate: | pool axis | composition | |---|---| | rows | 27,894 over 24,779 distinct targets | | corruption band | medium 37%, heavy 35%, light 24%, identity/no-op 4% | | `len_mode` | match 61%, inflate 29%, compress 10% | | kind | prose 97%, dialogue 2%, structural no-ops 1% | | target source | r/WritingPrompts 50%, AO3 39%, roleplay forum 9%, fanfiction.net 1% | Pairs pass invariant gates before they reach the GPU: POV, tense, who is in the scene, grammatical correctness on the target side, content recall stratified by target length, and NLI entailment both ways. Two further screens shape what reaches training — a floor on how much a pair actually changes, and a floor on the sentence-length variety of the target, applied at a higher threshold for long paragraphs than short ones. **Loss on the target paragraph only.** Everything before `rewrite` is masked. | | | |---|---| | LoRA | r=32, alpha=64, dropout 0.05 | | target modules | q, k, v, o, gate, up, down, **and `lm_head`** | | trainable | 71,004,160 params (1.73%) | | schedule | 2 epochs, lr 2e-4 cosine, batch 8 × accum 4, seq 2048 | | steps | 1,706 on one RTX 3090 | ## The merge Merged at strength 1.20. Rank 32 with alpha 64 is a LoRA scaling of 2.0, so the effective scaling is **2.40**: `W + (B @ A) * 2.40`. Merged in float32, stored bfloat16. `lm_head` is adapted, and `Qwen3-4B-Base` ties `lm_head.weight` to `embed_tokens.weight`. **This checkpoint is untied**: the merged output head is stored separately and the input embeddings are bit-identical to the base model's, which is what training assumed. `config.json` says `tie_word_embeddings: false` and it means it. Do not re-tie it, and if you convert to another format, check that the head survived. ## Limitations - **Not an instruct model.** It has one job and one prompt. There is nothing to ask it. - **Works on fictional prose only.** May not work on technical documentation. - **One paragraph per call.** Longer input degrades; split it. - **Will not pass AI detectors.** Pangram and such will still know because this model preserves word choices and certain sentence structures. - **English only**, narrative register (third and first person fiction, dialogue with quoted speech). - **Short input pads and invents.** The floor is about 15 words, and below it the failure is fabrication rather than gibberish. See [Input length](#input-length). - **Repeats a noun sooner than an LLM would.** Human prose reuses a plain noun where generated prose uses a synonym, and this model has learned that habit. ## License The weights in this repository are released under the **GNU Affero General Public License, version 3**. The full text is in `LICENSE`. This is a derivative of [`Qwen/Qwen3-4B-Base`](https://huggingface.co/Qwen/Qwen3-4B-Base), which is licensed under **Apache License 2.0**. That license is preserved and its terms continue to apply to the base weights this model was built from; the AGPL covers the combined work as distributed here. Apache-2.0 is one-way compatible with AGPLv3, which is what makes this combination possible. If you run a modified version of this model as a network service, AGPL section 13 requires you to offer the corresponding source of your modifications to its users.