How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "Sudhanshu1985/slm-125m-raft"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Sudhanshu1985/slm-125m-raft",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker
docker model run hf.co/Sudhanshu1985/slm-125m-raft
Quick Links

slm-125m-raft

A RAFT (Retrieval-Augmented Fine-Tuning) version of the 125M legal/financial model: thesreedath/slm-125m-base fine-tuned to answer from a retrieved set of documents (one golden chunk + distractors) and to refuse when the answer is not present in any of them.

Companion to Sudhanshu1985/slm-125m-sft (plain grounded-QA SFT). RAFT adds distractor documents and a 30% "no-golden" negative rate, so the model learns to ignore irrelevant retrieved chunks and not hallucinate.

Results (held-out validation)

Base RAFT
Loss 2.81 0.85
Perplexity 16.54 2.34

(Val set is the harder RAFT set โ€” multi-document with distractors.)

Training data

Sudhanshu1985/slm-125m-raft-dataset โ€” 23,830 examples derived from the reviewed QA set:

  • 70% positive: question + [golden chunk + 3 distractors] (shuffled) โ†’ grounded answer
  • 30% negative: golden removed / unanswerable โ†’ "That is not stated in the context."
  • ~200-token chunks so golden + distractors fit the 1024 window.

Recipe: 3 epochs, 1xH100, lr 2e-5 cosine, AdamW, bf16, effective batch 32, loss on the assistant span only.

Chat template

<|bos|><|system|>SYSTEM<|user|>USER<|assistant|>ANSWER<|eos|>

The user turn holds the retrieved documents followed by the question, e.g.:

Context:
[Document 1]
...
[Document 2]
...

<your question>
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Sudhanshu1985/slm-125m-raft")
model = AutoModelForCausalLM.from_pretrained("Sudhanshu1985/slm-125m-raft")

Limitations

  • Knows no facts of its own; works only over the documents you supply.
  • 1024-token window โ€” keep the retrieved doc set short.
  • Domain-biased toward US legal/financial register. Not legal or financial advice.
Downloads last month
15
Safetensors
Model size
0.1B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Sudhanshu1985/slm-125m-raft

Finetuned
(14)
this model

Dataset used to train Sudhanshu1985/slm-125m-raft