--- license: apache-2.0 base_model: Qwen/Qwen3.5-9B base_model_relation: finetune language: - en library_name: transformers pipeline_tag: text-generation datasets: - ThakiCloud/ChainCheck tags: - rag - evidence-sufficiency - multi-hop-qa - llm-judge - chaincheck - qwen --- # ChainCheck-Judge-Qwen3.5-9B **An evidence-sufficiency judge for RAG whose score tracks whether the reasoning chain is intact, not merely whether the passages were edited** (2Wiki chain selectivity +0.078 [-0.003, +0.156] → +0.386 [+0.329, +0.441] against the same-size zero-shot model). Given a question and retrieved passages, the model returns a log-odds score for "the passages together are sufficient to answer". It is Qwen/Qwen3.5-9B with a LoRA trained on ChainCheck counterfactual triplets, **merged back into the base as a single set of full bf16 weights** (the base's MTP tensors are carried over unchanged, so MTP speculative decoding still works). Why it matters: an ordinary sufficiency benchmark only asks whether a system separates untouched passages from broken ones, and a near-perfect score there can come from reacting to *any* edit. In a RAG pipeline that means an edited-but-valid context gets flagged and a subtly broken chain gets through. ChainCheck adds a matched control that separates the two. ## Measurements (2026-10-05, conditions stated) Same prompt, same scoring pipeline, one B200 GPU. Pairs a 27B model answers closed-book are excluded. Intervals are 95% pair-bootstrap (2,000 resamples). Σ = chain effect − |edit effect|; a lower bound above zero means the score reacts more to breaking the chain than to editing the passages. | Axis | zero-shot base | This model | |---|---|---| | 2Wiki nominal AUC, real entities (n=295 pairs) | 0.898 | 0.993 | | Chain selectivity Σ, MuSiQue test, real (n=324) | -0.071 [-0.154, +0.009] | +0.238 [+0.194, +0.284] | | Chain selectivity Σ, MuSiQue confirmation, real (n=122) | -0.016 [-0.147, +0.115] | +0.107 [+0.057, +0.164] | | Chain selectivity Σ, 2Wiki replication, real (n=295) | +0.078 [-0.003, +0.156] | +0.386 [+0.329, +0.441] | | Chain selectivity Σ, MuSiQue test, synthetic (n=224) | +0.134 [+0.036, +0.232] | +0.107 [+0.067, +0.147] | | Chain selectivity Σ, MuSiQue confirmation, synthetic (n=98) | +0.255 [+0.102, +0.357] | +0.041 [+0.010, +0.082] | | Chain selectivity Σ, 2Wiki replication, synthetic (n=141) | +0.192 [+0.092, +0.291] | +0.199 [+0.135, +0.270] | Honest caveats: - The tuned model tends to score the consistently edited context (B) above the untouched one (A): the edit effect is negative in every cell (range -0.459 to -0.107). Σ subtracts |EE|, so this is already penalised in the table, but the score is not edit-invariant. - With synthetic replacement entities, chain selectivity is lower after tuning than for the same-size zero-shot model on MuSiQue test (+0.134 → +0.107), MuSiQue confirmation (+0.255 → +0.041). The gains are on real entities; do not assume they transfer to entities the model has never seen. - 2Wiki was never used for training or model selection, and MuSiQue confirmation is a held-out slice built after the protocol was frozen. The release gate (2Wiki Σ lower bound > 0 on real and synthetic, MuSiQue test real Σ lower bound > 0, 2Wiki nominal no more than 0.02 below zero-shot) was fixed before any score was seen. ## Usage ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo = "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto") SYSTEM = ("You are a strict evidence auditor for a retrieval-augmented QA system. " "Decide whether the given passages, taken together, contain enough information " "to fully answer the question. Related-but-insufficient passages do NOT count.") USER = ("Question:\n{query}\n\nPassages:\n{passages}\n\n" "Do the passages together contain sufficient evidence to answer the question? " "Answer with a single word: yes or no.") yes = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"yes", "y"}] no = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"no", "n"}] def sufficiency_score(query, passages): body = "\n\n".join(f"[{i + 1}] {p[:4000]}" for i, p in enumerate(passages)) msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": USER.format(query=query, passages=body)}] text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False) ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device) with torch.no_grad(): lp = torch.log_softmax(model(**ids).logits[0, -1].float(), -1) return (torch.logsumexp(lp[yes], -1) - torch.logsumexp(lp[no], -1)).item() # > 0 leans "sufficient" ``` The tokenizer, vocabulary and architecture are identical to the base (standard load). The score is a log-odds, not a calibrated probability: choose the threshold on your own validation data and use it to gate generation (answer, retrieve more, or abstain). ## Limitations - English Wikipedia multi-hop questions only (MuSiQue, 2WikiMultiHopQA). The edits are entity substitutions on two-hop chains; numerical errors, temporal staleness and longer chains are not measured. - A positive Σ says the score reacts more to chain breaks than to edits on this benchmark. It does not certify that answers built on accepted contexts are correct. - Not a chat model. It is meant to be read through the yes/no logits shown above, not through free generation. ## Benchmark All three sizes, each next to its own zero-shot baseline: | Model | Set | Nominal (real) | Σ real [95% CI] | Σ synthetic [95% CI] | |---|---|---|---|---| | Qwen3.5-4B zero-shot | MuSiQue test | 0.889 | +0.028 [-0.052, +0.111] | +0.188 [+0.094, +0.277] | | Qwen3.5-4B zero-shot | MuSiQue confirmation | 0.885 | +0.090 [-0.049, +0.222] | +0.327 [+0.194, +0.378] | | Qwen3.5-4B zero-shot | 2Wiki replication | 0.932 | +0.136 [+0.058, +0.214] | +0.319 [+0.220, +0.411] | | Qwen3.5-4B ChainCheck-tuned | MuSiQue test | 0.982 | +0.185 [+0.145, +0.228] | +0.027 [+0.009, +0.049] | | Qwen3.5-4B ChainCheck-tuned | MuSiQue confirmation | 0.959 | +0.098 [+0.041, +0.156] | -0.010 [-0.031, +0.000] | | Qwen3.5-4B ChainCheck-tuned | 2Wiki replication | 0.983 | +0.376 [+0.322, +0.437] | +0.071 [+0.028, +0.114] | | Qwen3.5-9B zero-shot | MuSiQue test | 0.874 | -0.071 [-0.154, +0.009] | +0.134 [+0.036, +0.232] | | Qwen3.5-9B zero-shot | MuSiQue confirmation | 0.877 | -0.016 [-0.147, +0.115] | +0.255 [+0.102, +0.357] | | Qwen3.5-9B zero-shot | 2Wiki replication | 0.898 | +0.078 [-0.003, +0.156] | +0.192 [+0.092, +0.291] | | Qwen3.5-9B ChainCheck-tuned | MuSiQue test | 0.991 | +0.238 [+0.194, +0.284] | +0.107 [+0.067, +0.147] | | Qwen3.5-9B ChainCheck-tuned | MuSiQue confirmation | 1.000 | +0.107 [+0.057, +0.164] | +0.041 [+0.010, +0.082] | | Qwen3.5-9B ChainCheck-tuned | 2Wiki replication | 0.993 | +0.386 [+0.329, +0.441] | +0.199 [+0.135, +0.270] | | Qwen3.8-27B zero-shot | MuSiQue test | 0.895 | +0.201 [+0.123, +0.278] | +0.388 [+0.312, +0.420] | | Qwen3.8-27B zero-shot | MuSiQue confirmation | 0.943 | +0.180 [+0.066, +0.295] | +0.418 [+0.316, +0.469] | | Qwen3.8-27B zero-shot | 2Wiki replication | 0.905 | +0.129 [+0.054, +0.200] | +0.397 [+0.305, +0.454] | | Qwen3.8-27B ChainCheck-tuned | MuSiQue test | 1.000 | +0.225 [+0.182, +0.269] | +0.125 [+0.085, +0.165] | | Qwen3.8-27B ChainCheck-tuned | MuSiQue confirmation | 1.000 | +0.107 [+0.057, +0.164] | +0.051 [+0.010, +0.102] | | Qwen3.8-27B ChainCheck-tuned | 2Wiki replication | 0.997 | +0.461 [+0.397, +0.485] | +0.170 [+0.106, +0.234] | - Qwen3.5-4B: released - Qwen3.5-9B: released - Qwen3.8-27B: released Benchmark kit and standalone evaluator: [ThakiCloud/ChainCheck](https://huggingface.co/datasets/ThakiCloud/ChainCheck). ## Training LoRA (rank 8, α 16, attention q/k/v/o on full-attention layers, lr 1e-4, gradient accumulation 8, seed 0, one epoch over 7,481 MuSiQue train triplets) with labels A = 1, B = 1, D = 0 and loss = binary cross-entropy + max(0, 1 − (min(s_A, s_B) − s_D)). After merging, the full weights were reloaded from disk and re-scored on 64 held-out items to confirm they reproduce the adapter scores. ## License Apache-2.0 (base: Qwen/Qwen3.5-9B). Training and evaluation passages derive from MuSiQue (CC BY 4.0) and 2WikiMultiHopQA (Apache-2.0); passage text originates from Wikipedia. ## Paper ChainCheck: When Near-Perfect Evidence-Sufficiency Scores Fail to Distinguish Chain-Sensitive Systems — preprint forthcoming The paper shows that near-perfect nominal evidence scores do not reliably distinguish chain-sensitive systems, and that separating chain sensitivity from edit sensitivity changes the apparent ordering of systems. This model applies that protocol as a training signal. Read its gains against the same-size zero-shot column, not against the nominal score alone, which is close to its ceiling for every system.