Instructions to use ThakiCloud/ChainCheck-Judge-Qwen3.5-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ThakiCloud/ChainCheck-Judge-Qwen3.5-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ThakiCloud/ChainCheck-Judge-Qwen3.5-9B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ThakiCloud/ChainCheck-Judge-Qwen3.5-9B") model = AutoModelForMultimodalLM.from_pretrained("ThakiCloud/ChainCheck-Judge-Qwen3.5-9B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ThakiCloud/ChainCheck-Judge-Qwen3.5-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ThakiCloud/ChainCheck-Judge-Qwen3.5-9B
- SGLang
How to use ThakiCloud/ChainCheck-Judge-Qwen3.5-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ThakiCloud/ChainCheck-Judge-Qwen3.5-9B with Docker Model Runner:
docker model run hf.co/ThakiCloud/ChainCheck-Judge-Qwen3.5-9B
ChainCheck-Judge-Qwen3.5-9B
An evidence-sufficiency judge for RAG whose score tracks whether the reasoning chain is intact, not merely whether the passages were edited (2Wiki chain selectivity +0.078 [-0.003, +0.156] → +0.386 [+0.329, +0.441] against the same-size zero-shot model).
Given a question and retrieved passages, the model returns a log-odds score for "the passages together are sufficient to answer". It is Qwen/Qwen3.5-9B with a LoRA trained on ChainCheck counterfactual triplets, merged back into the base as a single set of full bf16 weights (the base's MTP tensors are carried over unchanged, so MTP speculative decoding still works).
Why it matters: an ordinary sufficiency benchmark only asks whether a system separates untouched passages from broken ones, and a near-perfect score there can come from reacting to any edit. In a RAG pipeline that means an edited-but-valid context gets flagged and a subtly broken chain gets through. ChainCheck adds a matched control that separates the two.
Measurements (2026-10-05, conditions stated)
Same prompt, same scoring pipeline, one B200 GPU. Pairs a 27B model answers closed-book are excluded. Intervals are 95% pair-bootstrap (2,000 resamples). Σ = chain effect − |edit effect|; a lower bound above zero means the score reacts more to breaking the chain than to editing the passages.
| Axis | zero-shot base | This model |
|---|---|---|
| 2Wiki nominal AUC, real entities (n=295 pairs) | 0.898 | 0.993 |
| Chain selectivity Σ, MuSiQue test, real (n=324) | -0.071 [-0.154, +0.009] | +0.238 [+0.194, +0.284] |
| Chain selectivity Σ, MuSiQue confirmation, real (n=122) | -0.016 [-0.147, +0.115] | +0.107 [+0.057, +0.164] |
| Chain selectivity Σ, 2Wiki replication, real (n=295) | +0.078 [-0.003, +0.156] | +0.386 [+0.329, +0.441] |
| Chain selectivity Σ, MuSiQue test, synthetic (n=224) | +0.134 [+0.036, +0.232] | +0.107 [+0.067, +0.147] |
| Chain selectivity Σ, MuSiQue confirmation, synthetic (n=98) | +0.255 [+0.102, +0.357] | +0.041 [+0.010, +0.082] |
| Chain selectivity Σ, 2Wiki replication, synthetic (n=141) | +0.192 [+0.092, +0.291] | +0.199 [+0.135, +0.270] |
Honest caveats:
- The tuned model tends to score the consistently edited context (B) above the untouched one (A): the edit effect is negative in every cell (range -0.459 to -0.107). Σ subtracts |EE|, so this is already penalised in the table, but the score is not edit-invariant.
- With synthetic replacement entities, chain selectivity is lower after tuning than for the same-size zero-shot model on MuSiQue test (+0.134 → +0.107), MuSiQue confirmation (+0.255 → +0.041). The gains are on real entities; do not assume they transfer to entities the model has never seen.
- 2Wiki was never used for training or model selection, and MuSiQue confirmation is a held-out slice built after the protocol was frozen. The release gate (2Wiki Σ lower bound > 0 on real and synthetic, MuSiQue test real Σ lower bound > 0, 2Wiki nominal no more than 0.02 below zero-shot) was fixed before any score was seen.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "ThakiCloud/ChainCheck-Judge-Qwen3.5-9B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto")
SYSTEM = ("You are a strict evidence auditor for a retrieval-augmented QA system. "
"Decide whether the given passages, taken together, contain enough information "
"to fully answer the question. Related-but-insufficient passages do NOT count.")
USER = ("Question:\n{query}\n\nPassages:\n{passages}\n\n"
"Do the passages together contain sufficient evidence to answer the question? "
"Answer with a single word: yes or no.")
yes = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"yes", "y"}]
no = [i for i in range(len(tok)) if tok.decode([i]).strip().lower() in {"no", "n"}]
def sufficiency_score(query, passages):
body = "\n\n".join(f"[{i + 1}] {p[:4000]}" for i, p in enumerate(passages))
msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": USER.format(query=query, passages=body)}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False)
ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
lp = torch.log_softmax(model(**ids).logits[0, -1].float(), -1)
return (torch.logsumexp(lp[yes], -1) - torch.logsumexp(lp[no], -1)).item() # > 0 leans "sufficient"
The tokenizer, vocabulary and architecture are identical to the base (standard load). The score is a log-odds, not a calibrated probability: choose the threshold on your own validation data and use it to gate generation (answer, retrieve more, or abstain).
Limitations
- English Wikipedia multi-hop questions only (MuSiQue, 2WikiMultiHopQA). The edits are entity substitutions on two-hop chains; numerical errors, temporal staleness and longer chains are not measured.
- A positive Σ says the score reacts more to chain breaks than to edits on this benchmark. It does not certify that answers built on accepted contexts are correct.
- Not a chat model. It is meant to be read through the yes/no logits shown above, not through free generation.
Benchmark
All three sizes, each next to its own zero-shot baseline:
| Model | Set | Nominal (real) | Σ real [95% CI] | Σ synthetic [95% CI] |
|---|---|---|---|---|
| Qwen3.5-4B zero-shot | MuSiQue test | 0.889 | +0.028 [-0.052, +0.111] | +0.188 [+0.094, +0.277] |
| Qwen3.5-4B zero-shot | MuSiQue confirmation | 0.885 | +0.090 [-0.049, +0.222] | +0.327 [+0.194, +0.378] |
| Qwen3.5-4B zero-shot | 2Wiki replication | 0.932 | +0.136 [+0.058, +0.214] | +0.319 [+0.220, +0.411] |
| Qwen3.5-4B ChainCheck-tuned | MuSiQue test | 0.982 | +0.185 [+0.145, +0.228] | +0.027 [+0.009, +0.049] |
| Qwen3.5-4B ChainCheck-tuned | MuSiQue confirmation | 0.959 | +0.098 [+0.041, +0.156] | -0.010 [-0.031, +0.000] |
| Qwen3.5-4B ChainCheck-tuned | 2Wiki replication | 0.983 | +0.376 [+0.322, +0.437] | +0.071 [+0.028, +0.114] |
| Qwen3.5-9B zero-shot | MuSiQue test | 0.874 | -0.071 [-0.154, +0.009] | +0.134 [+0.036, +0.232] |
| Qwen3.5-9B zero-shot | MuSiQue confirmation | 0.877 | -0.016 [-0.147, +0.115] | +0.255 [+0.102, +0.357] |
| Qwen3.5-9B zero-shot | 2Wiki replication | 0.898 | +0.078 [-0.003, +0.156] | +0.192 [+0.092, +0.291] |
| Qwen3.5-9B ChainCheck-tuned | MuSiQue test | 0.991 | +0.238 [+0.194, +0.284] | +0.107 [+0.067, +0.147] |
| Qwen3.5-9B ChainCheck-tuned | MuSiQue confirmation | 1.000 | +0.107 [+0.057, +0.164] | +0.041 [+0.010, +0.082] |
| Qwen3.5-9B ChainCheck-tuned | 2Wiki replication | 0.993 | +0.386 [+0.329, +0.441] | +0.199 [+0.135, +0.270] |
| Qwen3.8-27B zero-shot | MuSiQue test | 0.895 | +0.201 [+0.123, +0.278] | +0.388 [+0.312, +0.420] |
| Qwen3.8-27B zero-shot | MuSiQue confirmation | 0.943 | +0.180 [+0.066, +0.295] | +0.418 [+0.316, +0.469] |
| Qwen3.8-27B zero-shot | 2Wiki replication | 0.905 | +0.129 [+0.054, +0.200] | +0.397 [+0.305, +0.454] |
| Qwen3.8-27B ChainCheck-tuned | MuSiQue test | 1.000 | +0.225 [+0.182, +0.269] | +0.125 [+0.085, +0.165] |
| Qwen3.8-27B ChainCheck-tuned | MuSiQue confirmation | 1.000 | +0.107 [+0.057, +0.164] | +0.051 [+0.010, +0.102] |
| Qwen3.8-27B ChainCheck-tuned | 2Wiki replication | 0.997 | +0.461 [+0.397, +0.485] | +0.170 [+0.106, +0.234] |
- Qwen3.5-4B: released
- Qwen3.5-9B: released
- Qwen3.8-27B: released
Benchmark kit and standalone evaluator: ThakiCloud/ChainCheck.
Training
LoRA (rank 8, α 16, attention q/k/v/o on full-attention layers, lr 1e-4, gradient accumulation 8, seed 0, one epoch over 7,481 MuSiQue train triplets) with labels A = 1, B = 1, D = 0 and loss = binary cross-entropy + max(0, 1 − (min(s_A, s_B) − s_D)). After merging, the full weights were reloaded from disk and re-scored on 64 held-out items to confirm they reproduce the adapter scores.
License
Apache-2.0 (base: Qwen/Qwen3.5-9B). Training and evaluation passages derive from MuSiQue (CC BY 4.0) and 2WikiMultiHopQA (Apache-2.0); passage text originates from Wikipedia.
Paper
ChainCheck: When Near-Perfect Evidence-Sufficiency Scores Fail to Distinguish Chain-Sensitive Systems — preprint forthcoming
The paper shows that near-perfect nominal evidence scores do not reliably distinguish chain-sensitive systems, and that separating chain sensitivity from edit sensitivity changes the apparent ordering of systems. This model applies that protocol as a training signal. Read its gains against the same-size zero-shot column, not against the nominal score alone, which is close to its ceiling for every system.
- Downloads last month
- 330