--- license: apache-2.0 base_model: Qwen/Qwen3.5-9B base_model_relation: finetune library_name: transformers pipeline_tag: text-generation language: [en] datasets: [ThakiCloud/ChainCheck] tags: [rag, evidence-sufficiency, multi-hop-qa, judge, chaincheck] --- # ChainCheck-Judge-Qwen3.5-9B An evidence-sufficiency judge for retrieval-augmented QA: given a question and retrieved passages, it scores whether the passages together are enough to answer. It is Qwen/Qwen3.5-9B fine-tuned so that its score follows **whether the reasoning chain is intact**, not merely whether the passages were edited. Why this matters: ordinary sufficiency benchmarks reward a system for separating untouched passages from broken ones, and near-perfect scores there can come from reacting to *any* edit. In a RAG pipeline that means an edited-but-valid context gets flagged and a subtly broken chain gets through. This model is trained and evaluated with the ChainCheck counterfactual controls described in *ChainCheck: When Near-Perfect Evidence-Sufficiency Scores Fail to Distinguish Chain-Sensitive Systems* (preprint forthcoming). ## Usage ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo = "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto") SYSTEM = ("You are a strict evidence auditor for a retrieval-augmented QA system. " "Decide whether the given passages, taken together, contain enough information " "to fully answer the question. Related-but-insufficient passages do NOT count.") USER = ("Question:\n{query}\n\nPassages:\n{passages}\n\n" "Do the passages together contain sufficient evidence to answer the question? " "Answer with a single word: yes or no.") yes = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"yes", "y"}] no = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"no", "n"}] def sufficiency_score(query, passages): body = "\n\n".join(f"[{i + 1}] {p[:4000]}" for i, p in enumerate(passages)) msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": USER.format(query=query, passages=body)}] text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False) ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device) with torch.no_grad(): lp = torch.log_softmax(model(**ids).logits[0, -1].float(), -1) return (torch.logsumexp(lp[yes], -1) - torch.logsumexp(lp[no], -1)).item() # > 0 leans "sufficient" ``` Use the score to gate generation (answer, retrieve more, or abstain). Pick the threshold on your own validation data; the score is a log-odds, not a calibrated probability. ## Evaluation ChainCheck builds three cells per question: **A** (chain intact, no edit), **B** (chain intact, consistent entity swap), **D** (chain broken by a later-hop substitution). Nominal = AUC(A > D) is what an ordinary benchmark reports. Chain effect CE = AUC(B > D) − 0.5, edit effect EE = AUC(A > B) − 0.5, and chain selectivity **Σ = CE − |EE|**. Σ > 0 with its 95% lower bound above zero means the score tracks the chain more than the edit. Intervals are pair-bootstrap (2,000 resamples). Pairs a 27B model answers closed-book are excluded. "Synthetic" replaces the bridge entity with a fictional one. All three sizes, with the same-size zero-shot model as baseline (same prompt, no training): | Model | Set | Nominal (real) | Σ real [95% CI] | Σ synthetic [95% CI] | |---|---|---|---|---| | Qwen3.5-4B zero-shot | MuSiQue test | 0.889 | +0.028 [-0.052, +0.111] | +0.188 [+0.094, +0.277] | | Qwen3.5-4B zero-shot | MuSiQue confirmation | 0.885 | +0.090 [-0.049, +0.222] | +0.327 [+0.194, +0.378] | | Qwen3.5-4B zero-shot | 2Wiki replication | 0.932 | +0.136 [+0.058, +0.214] | +0.319 [+0.220, +0.411] | | Qwen3.5-4B ChainCheck-tuned | MuSiQue test | 0.982 | +0.185 [+0.145, +0.228] | +0.027 [+0.009, +0.049] | | Qwen3.5-4B ChainCheck-tuned | MuSiQue confirmation | 0.959 | +0.098 [+0.041, +0.156] | -0.010 [-0.031, +0.000] | | Qwen3.5-4B ChainCheck-tuned | 2Wiki replication | 0.983 | +0.376 [+0.322, +0.437] | +0.071 [+0.028, +0.114] | | Qwen3.5-9B zero-shot | MuSiQue test | 0.874 | -0.071 [-0.154, +0.009] | +0.134 [+0.036, +0.232] | | Qwen3.5-9B zero-shot | MuSiQue confirmation | 0.877 | -0.016 [-0.147, +0.115] | +0.255 [+0.102, +0.357] | | Qwen3.5-9B zero-shot | 2Wiki replication | 0.898 | +0.078 [-0.003, +0.156] | +0.192 [+0.092, +0.291] | | Qwen3.5-9B ChainCheck-tuned | MuSiQue test | 0.991 | +0.238 [+0.194, +0.284] | +0.107 [+0.067, +0.147] | | Qwen3.5-9B ChainCheck-tuned | MuSiQue confirmation | 1.000 | +0.107 [+0.057, +0.164] | +0.041 [+0.010, +0.082] | | Qwen3.5-9B ChainCheck-tuned | 2Wiki replication | 0.993 | +0.386 [+0.329, +0.441] | +0.199 [+0.135, +0.270] | | Qwen3.8-27B zero-shot | MuSiQue test | 0.895 | +0.201 [+0.123, +0.278] | +0.388 [+0.312, +0.420] | | Qwen3.8-27B zero-shot | MuSiQue confirmation | 0.943 | +0.180 [+0.066, +0.295] | +0.418 [+0.316, +0.469] | | Qwen3.8-27B zero-shot | 2Wiki replication | 0.905 | +0.129 [+0.054, +0.200] | +0.397 [+0.305, +0.454] | | Qwen3.8-27B ChainCheck-tuned | MuSiQue test | 1.000 | +0.225 [+0.182, +0.269] | +0.125 [+0.085, +0.165] | | Qwen3.8-27B ChainCheck-tuned | MuSiQue confirmation | 1.000 | +0.107 [+0.057, +0.164] | +0.051 [+0.010, +0.102] | | Qwen3.8-27B ChainCheck-tuned | 2Wiki replication | 0.997 | +0.461 [+0.397, +0.485] | +0.170 [+0.106, +0.234] | Every row is scored by the same release pipeline, so tuned and zero-shot rows are directly comparable. The Qwen3.8-27B zero-shot row is the paper's reference judge re-scored this way; it ranks items essentially identically to the paper's run, but A and B differ by one entity and score almost the same, so small bf16 numerical differences flip some A-versus-B comparisons and Σ moves within the paper's reported interval. Release gate, fixed before any score was seen: 2Wiki Σ lower bound > 0 on both real and synthetic, MuSiQue test real Σ lower bound > 0, and 2Wiki real nominal no more than 0.02 below the same-size zero-shot model. - Qwen3.5-4B: released - Qwen3.5-9B: released - Qwen3.8-27B: released 2Wiki was never used for training or model selection. MuSiQue confirmation is a held-out slice built after the protocol was frozen. ## Training LoRA (rank 8, α 16, attention q/k/v/o on full-attention layers, lr 1e-4, gradient accumulation 8, seed 0, one epoch) on ChainCheck MuSiQue train triplets (A = 1, B = 1, D = 0), binary cross-entropy plus a margin term min(s_A, s_B) − s_D ≥ 1, then merged into full weights. The merge was checked by re-scoring a held-out sample and comparing with the adapter scores. ## Limitations The tuned model scores the consistently edited context (B) above the untouched one (A): the edit effect is negative in every cell (-0.459 to -0.107). Σ subtracts |EE|, so this is already penalised in the table, but it means the score is not edit-invariant. With synthetic replacement entities, chain selectivity is lower after tuning than for the same-size zero-shot model on MuSiQue test (+0.134 → +0.107), MuSiQue confirmation (+0.255 → +0.041). The gains are on real entities; do not assume they transfer to entities the model has never seen. English Wikipedia multi-hop questions only (MuSiQue, 2WikiMultiHopQA). The edits are entity substitutions on two-hop chains; other failure types (numerical errors, temporal staleness, long chains) are not measured. A positive Σ says the score reacts more to chain breaks than to edits on this benchmark; it does not certify that answers built on accepted contexts are correct. Rankings in the table are point estimates and some intervals overlap. ## Citation Han, H. (2026). ChainCheck: When Near-Perfect Evidence-Sufficiency Scores Fail to Distinguish Chain-Sensitive Systems. ThakiCloud.