--- license: agpl-3.0 base_model: Qwen/Qwen3-1.7B-Base pipeline_tag: text-generation language: - en tags: - prose - rewriting - style-transfer - creative-writing - deslop - qwen3 library_name: transformers --- # prose-rewriter-1.7b-v1.4 A paragraph-level **prose rewriter**: it takes prose written by a large model and re-renders it to be more human, preserving the semantics it was given. `Qwen/Qwen3-1.7B-Base` with a rank-32 LoRA merged in at strength 1.025. Successor to [prose-rewriter-1.7b-v1.2](https://huggingface.co/chartreuse-verte/prose-rewriter-1.7b-v1.2). It **restructures more and copies less**, trained on more sentence-length rows. See [Evaluation](#evaluation). A larger version: [4B](https://huggingface.co/chartreuse-verte/prose-rewriter-4b-v1.3). ## Variants | Path | Format | Use with | |---|---|---| | `/` | safetensors bf16, `qwen3` arch | transformers | | `GGUF/prose-rewriter-1.7b-v1.4-Q8_0.gguf` | GGUF Q8_0, 2.17 GB | llama.cpp / llama-cpp-python | | `GGUF/prose-rewriter-1.7b-v1.4-Q4_K_M.gguf` | GGUF Q4_K_M, 1.28 GB | llama.cpp / llama-cpp-python | The quants carry the chat template and stop on `<|im_end|>`, and the adapted output head is kept separate from the token embeddings in both — Q8_0 stores it at Q8_0, Q4_K_M at Q6_K. ## Prompt format ``` <|im_start|>source {paragraph}<|im_end|> <|im_start|>edit match<|im_end|> <|im_start|>rewrite ``` The chat template in this repo builds exactly that string, byte for byte, from two roles: ```python messages = [ {"role": "source", "content": paragraph}, {"role": "edit", "content": "match"}, ] tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) ``` It is not a chat model. The template rejects an `edit` value outside the three modes rather than quietly building a prompt the weights have never seen. Any other role is treated as the source paragraph, so a runtime that probes the template with a `user` message still gets a valid prompt. ## The `edit` block is mandatory `edit` names which of three length transforms is being asked for. **The values describe the input, not the instruction.** They say what kind of text you are handing over: | `edit` | what it says about the input | what the model does | |---|---|---| | `match` | the source is about the length it should be | rewrite in place | | `inflate` | the source is padded relative to what it should be | cut | | `compress` | the source is flattened and too short | open it back out | `match` is the setting for "rewrite it, do not trim it". It's strongly recommended you use this mode. Sending **no** block is the worst thing you can do to this checkpoint. It was trained with the block, so omitting it collapses the model onto its deletion-heaviest mode. ## Serving recipe Sampled at `temperature=0.9, top_p=0.9`. Temperature 0.9 has been tested and internally to be the most optimal value. It's recommended you use this. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo = "chartreuse-verte/prose-rewriter-1.7b-v1.4" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="cuda").eval() def rewrite(paragraph, mode="match"): text = tok.apply_chat_template( [{"role": "source", "content": paragraph}, {"role": "edit", "content": mode}], tokenize=False, add_generation_prompt=True, ) ids = tok(text, return_tensors="pt", add_special_tokens=False).input_ids.to(model.device) out = model.generate(ids, max_new_tokens=512, do_sample=True, temperature=0.9, top_p=0.9) return tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True).strip() ``` Same thing under llama.cpp. The roles are `source` and `edit`, which no chat API models, so build the string yourself; `<|im_end|>` stops it: ```bash llama-cli -m GGUF/prose-rewriter-1.7b-v1.4-Q8_0.gguf -no-cnv -n 512 --temp 0.9 --top-p 0.9 \ -p '<|im_start|>source {paragraph}<|im_end|> <|im_start|>edit match<|im_end|> <|im_start|>rewrite ' ``` ## Input length The training pool's median input is 39 words and most of it is under 70, so serve it on anything from a full sentence up. The practical floor is about **15 words**. Below it the failure mode is padding and fabrication rather than gibberish: the model stretches the line toward its learned length and adds material the input never supported. Below 80 bytes, pass the text through unchanged. ## Evaluation v1.2 against v1.4 on 1,095 held-out paragraphs of LLM-written prose that neither model saw in training, both sent the identical prompt at `temperature=0.9, top_p=0.9`. Paired over the 1,068 inputs of 51 words or more: | | v1.2 | v1.4 | paired *t* | |---|---|---|---| | sentence-length variety vs input | +0.073 | **+0.127** | +7.52 | | words changed | 33.8% | **38.0%** | +6.37 | | words kept from the input | 0.707 | **0.679** | −5.70 | | length preserved | 0.881 | **0.905** | +4.90 | | sentence count moved | 71.4% | **77.0%** | +2.81 | | near-verbatim outputs | 3.9% | **3.3%** | −1.31 | | repeated 3-grams | 0.004 | 0.005 | +1.95 | **v1.4 rewrites more of the paragraph, keeps less of the original wording, and varies its sentence lengths substantially more, while holding on to more of the source's length.** Sentence-length variety is the largest single move and is the one most visible when reading: v1.2 tends to produce evenly-sized sentences, v1.4 mixes short and long the way human prose does. The one number pointing the other way is repeated 3-grams, up by 0.001 — heavier rewriting brings slightly more phrase repetition with it. ## Training **Corrupt forward, train backward.** The human paragraph is the target; an on-policy LLM manufactures the input by slop-ifying it. The target side is human prose: roughly 55/43 r/WritingPrompts ([`Mollymo/Human-to-AI-writing`](https://huggingface.co/datasets/Mollymo/Human-to-AI-writing)) and AO3 ([`midwestern-simulation-active/ao3_random_subset`](https://huggingface.co/datasets/midwestern-simulation-active/ao3_random_subset)), with a sliver of fanfiction.net ([`atom-in-the-universe/fanfics-10k-10k`](https://huggingface.co/datasets/atom-in-the-universe/fanfics-10k-10k)). The input side was generated by eleven corruptor endpoints — a few of which are two routes onto one set of weights — weighted and share-capped so no single model's tics dominate: | pool axis | composition | |---|---| | rows | 24,061 over 19,872 distinct targets | | corruptor | ds-flash-nano 17%, qwen-flash 15%, artemis-local 13%, gemma-local 11%, ds-flash 11%, gemma-nano 8%, qwen-local 7%, ox-alpha 5% + 4%, then artemis, muse, ds-pro | | corruption band | medium 37%, heavy 33%, light 26% | | `len_mode` | match 57%, inflate 28%, compress 10%, unmarked 5% | | kind | prose 97%, dialogue 2%, structural no-ops 1% | Pairs pass invariant gates before they reach the GPU: POV, tense, who is in the scene, grammatical correctness on the target side, content recall stratified by target length, and NLI entailment both ways. Two further screens shape what reaches training — a floor on how much a pair actually changes, and a floor on the sentence-length variety of the target, applied at a higher threshold for long paragraphs than short ones. **Loss on the target paragraph only.** Everything before `rewrite` is masked. | | | |---|---| | LoRA | r=32, alpha=64, dropout 0.05 | | target modules | q, k, v, o, gate, up, down, **and `lm_head`** | | trainable | 39,792,640 params (2.26%) | | schedule | 2 epochs, lr 1e-4 cosine, batch 4 × accum 8, seq 2048 | | steps | 1,468 on one RTX 3090, 50 min | | loss | train 1.070, val 1.055 (600 val rows, document-disjoint) | ## The merge Merged at strength 1.025. Rank 32 with alpha 64 is a LoRA scaling of 2.0, so the effective scaling is **2.05**: `W + (B @ A) * 2.05`. Merged in float32, stored bfloat16. `lm_head` is adapted, and `Qwen3-1.7B-Base` ties `lm_head.weight` to `embed_tokens.weight`. **This checkpoint is untied**: the merged output head is stored separately and the input embeddings are bit-identical to the base model's, which is what training assumed. `config.json` says `tie_word_embeddings: false` and it means it. Do not re-tie it, and if you convert to another format, check that the head survived. ## Limitations - **Not an instruct model.** It has one job and one prompt. There is nothing to ask it. - **Works on fictional prose only.** May not work on technical documentation. - **One paragraph per call.** Longer input degrades; split it. - **Will not pass AI detectors.** Pangram and such will still know because this model preserves word choices and certain sentence structures. - **English only**, narrative register (third and first person fiction, dialogue with quoted speech). - **Short input pads and invents.** The floor is about 15 words, and below it the failure is fabrication rather than gibberish. See [Input length](#input-length). - **Repeats a noun sooner than an LLM would.** Human prose reuses a plain noun where generated prose uses a synonym, and this model has learned that habit. ## License The weights in this repository are released under the **GNU Affero General Public License, version 3**. The full text is in `LICENSE`. This is a derivative of [`Qwen/Qwen3-1.7B-Base`](https://huggingface.co/Qwen/Qwen3-1.7B-Base), which is licensed under **Apache License 2.0**. That license is preserved and its terms continue to apply to the base weights this model was built from; the AGPL covers the combined work as distributed here. Apache-2.0 is one-way compatible with AGPLv3, which is what makes this combination possible. If you run a modified version of this model as a network service, AGPL section 13 requires you to offer the corresponding source of your modifications to its users.