ChainCheck-Judge-Qwen3.5-4B

An evidence-sufficiency judge for RAG whose score tracks whether the reasoning chain is intact, not merely whether the passages were edited (2Wiki chain selectivity +0.136 [+0.058, +0.214] → +0.376 [+0.322, +0.437] against the same-size zero-shot model).

Given a question and retrieved passages, the model returns a log-odds score for "the passages together are sufficient to answer". It is Qwen/Qwen3.5-4B with a LoRA trained on ChainCheck counterfactual triplets, merged back into the base as a single set of full bf16 weights (the base's MTP tensors are carried over unchanged, so MTP speculative decoding still works).

Why it matters: an ordinary sufficiency benchmark only asks whether a system separates untouched passages from broken ones, and a near-perfect score there can come from reacting to any edit. In a RAG pipeline that means an edited-but-valid context gets flagged and a subtly broken chain gets through. ChainCheck adds a matched control that separates the two.

Measurements (2026-10-05, conditions stated)

Same prompt, same scoring pipeline, one B200 GPU. Pairs a 27B model answers closed-book are excluded. Intervals are 95% pair-bootstrap (2,000 resamples). Σ = chain effect − |edit effect|; a lower bound above zero means the score reacts more to breaking the chain than to editing the passages.

Axis zero-shot base This model
2Wiki nominal AUC, real entities (n=295 pairs) 0.932 0.983
Chain selectivity Σ, MuSiQue test, real (n=324) +0.028 [-0.052, +0.111] +0.185 [+0.145, +0.228]
Chain selectivity Σ, MuSiQue confirmation, real (n=122) +0.090 [-0.049, +0.222] +0.098 [+0.041, +0.156]
Chain selectivity Σ, 2Wiki replication, real (n=295) +0.136 [+0.058, +0.214] +0.376 [+0.322, +0.437]
Chain selectivity Σ, MuSiQue test, synthetic (n=224) +0.188 [+0.094, +0.277] +0.027 [+0.009, +0.049]
Chain selectivity Σ, MuSiQue confirmation, synthetic (n=98) +0.327 [+0.194, +0.378] -0.010 [-0.031, +0.000]
Chain selectivity Σ, 2Wiki replication, synthetic (n=141) +0.319 [+0.220, +0.411] +0.071 [+0.028, +0.114]

Honest caveats:

  • The tuned model tends to score the consistently edited context (B) above the untouched one (A): the edit effect is negative in every cell (range -0.500 to -0.114). Σ subtracts |EE|, so this is already penalised in the table, but the score is not edit-invariant.
  • With synthetic replacement entities, chain selectivity is lower after tuning than for the same-size zero-shot model on MuSiQue test (+0.188 → +0.027), MuSiQue confirmation (+0.327 → -0.010), 2Wiki replication (+0.319 → +0.071). The gains are on real entities; do not assume they transfer to entities the model has never seen.
  • 2Wiki was never used for training or model selection, and MuSiQue confirmation is a held-out slice built after the protocol was frozen. The release gate (2Wiki Σ lower bound > 0 on real and synthetic, MuSiQue test real Σ lower bound > 0, 2Wiki nominal no more than 0.02 below zero-shot) was fixed before any score was seen.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ThakiCloud/ChainCheck-Judge-Qwen3.5-4B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto")

SYSTEM = ("You are a strict evidence auditor for a retrieval-augmented QA system. "
          "Decide whether the given passages, taken together, contain enough information "
          "to fully answer the question. Related-but-insufficient passages do NOT count.")
USER = ("Question:\n{query}\n\nPassages:\n{passages}\n\n"
        "Do the passages together contain sufficient evidence to answer the question? "
        "Answer with a single word: yes or no.")

yes = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"yes", "y"}]
no = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"no", "n"}]

def sufficiency_score(query, passages):
    body = "\n\n".join(f"[{i + 1}] {p[:4000]}" for i, p in enumerate(passages))
    msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": USER.format(query=query, passages=body)}]
    text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False)
    ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
    with torch.no_grad():
        lp = torch.log_softmax(model(**ids).logits[0, -1].float(), -1)
    return (torch.logsumexp(lp[yes], -1) - torch.logsumexp(lp[no], -1)).item()  # > 0 leans "sufficient"

The tokenizer, vocabulary and architecture are identical to the base (standard load). The score is a log-odds, not a calibrated probability: choose the threshold on your own validation data and use it to gate generation (answer, retrieve more, or abstain).

Limitations

  • English Wikipedia multi-hop questions only (MuSiQue, 2WikiMultiHopQA). The edits are entity substitutions on two-hop chains; numerical errors, temporal staleness and longer chains are not measured.
  • A positive Σ says the score reacts more to chain breaks than to edits on this benchmark. It does not certify that answers built on accepted contexts are correct.
  • Not a chat model. It is meant to be read through the yes/no logits shown above, not through free generation.

Benchmark

All three sizes, each next to its own zero-shot baseline:

Model Set Nominal (real) Σ real [95% CI] Σ synthetic [95% CI]
Qwen3.5-4B zero-shot MuSiQue test 0.889 +0.028 [-0.052, +0.111] +0.188 [+0.094, +0.277]
Qwen3.5-4B zero-shot MuSiQue confirmation 0.885 +0.090 [-0.049, +0.222] +0.327 [+0.194, +0.378]
Qwen3.5-4B zero-shot 2Wiki replication 0.932 +0.136 [+0.058, +0.214] +0.319 [+0.220, +0.411]
Qwen3.5-4B ChainCheck-tuned MuSiQue test 0.982 +0.185 [+0.145, +0.228] +0.027 [+0.009, +0.049]
Qwen3.5-4B ChainCheck-tuned MuSiQue confirmation 0.959 +0.098 [+0.041, +0.156] -0.010 [-0.031, +0.000]
Qwen3.5-4B ChainCheck-tuned 2Wiki replication 0.983 +0.376 [+0.322, +0.437] +0.071 [+0.028, +0.114]
Qwen3.5-9B zero-shot MuSiQue test 0.874 -0.071 [-0.154, +0.009] +0.134 [+0.036, +0.232]
Qwen3.5-9B zero-shot MuSiQue confirmation 0.877 -0.016 [-0.147, +0.115] +0.255 [+0.102, +0.357]
Qwen3.5-9B zero-shot 2Wiki replication 0.898 +0.078 [-0.003, +0.156] +0.192 [+0.092, +0.291]
Qwen3.5-9B ChainCheck-tuned MuSiQue test 0.991 +0.238 [+0.194, +0.284] +0.107 [+0.067, +0.147]
Qwen3.5-9B ChainCheck-tuned MuSiQue confirmation 1.000 +0.107 [+0.057, +0.164] +0.041 [+0.010, +0.082]
Qwen3.5-9B ChainCheck-tuned 2Wiki replication 0.993 +0.386 [+0.329, +0.441] +0.199 [+0.135, +0.270]
Qwen3.8-27B zero-shot MuSiQue test 0.895 +0.201 [+0.123, +0.278] +0.388 [+0.312, +0.420]
Qwen3.8-27B zero-shot MuSiQue confirmation 0.943 +0.180 [+0.066, +0.295] +0.418 [+0.316, +0.469]
Qwen3.8-27B zero-shot 2Wiki replication 0.905 +0.129 [+0.054, +0.200] +0.397 [+0.305, +0.454]
Qwen3.8-27B ChainCheck-tuned MuSiQue test 1.000 +0.225 [+0.182, +0.269] +0.125 [+0.085, +0.165]
Qwen3.8-27B ChainCheck-tuned MuSiQue confirmation 1.000 +0.107 [+0.057, +0.164] +0.051 [+0.010, +0.102]
Qwen3.8-27B ChainCheck-tuned 2Wiki replication 0.997 +0.461 [+0.397, +0.485] +0.170 [+0.106, +0.234]
  • Qwen3.5-4B: released
  • Qwen3.5-9B: released
  • Qwen3.8-27B: released

Benchmark kit and standalone evaluator: ThakiCloud/ChainCheck.

Training

LoRA (rank 8, α 16, attention q/k/v/o on full-attention layers, lr 1e-4, gradient accumulation 8, seed 0, one epoch over 7,481 MuSiQue train triplets) with labels A = 1, B = 1, D = 0 and loss = binary cross-entropy + max(0, 1 − (min(s_A, s_B) − s_D)). After merging, the full weights were reloaded from disk and re-scored on 64 held-out items to confirm they reproduce the adapter scores.

License

Apache-2.0 (base: Qwen/Qwen3.5-4B). Training and evaluation passages derive from MuSiQue (CC BY 4.0) and 2WikiMultiHopQA (Apache-2.0); passage text originates from Wikipedia.

Paper

ChainCheck: When Near-Perfect Evidence-Sufficiency Scores Fail to Distinguish Chain-Sensitive Systems — preprint forthcoming

The paper shows that near-perfect nominal evidence scores do not reliably distinguish chain-sensitive systems, and that separating chain sensitivity from edit sensitivity changes the apparent ordering of systems. This model applies that protocol as a training signal. Read its gains against the same-size zero-shot column, not against the nominal score alone, which is close to its ceiling for every system.

Downloads last month
322
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ThakiCloud/ChainCheck-Judge-Qwen3.5-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(885)
this model

Dataset used to train ThakiCloud/ChainCheck-Judge-Qwen3.5-4B

Collection including ThakiCloud/ChainCheck-Judge-Qwen3.5-4B