Kanha Qwen3 PIT checkpoint

Run identity

  • Run ID: kanha.ai-1.7b-pit-v1
  • Base model: Qwen/Qwen3-1.7B
  • Base model revision: 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e
  • Tokenizer revision: 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e
  • Training method: PIT document continuation plus Q&A
  • Final merged dtype: bfloat16
  • Source site: https://kanha.ai
  • Training identity hash: 5c52a30c64f842deaa6bc43e8d7f9960473112cff7eb39dcfe6044b86503ac1c
  • Train documents: 17
  • Evaluation documents: 0
  • Train Q&A pairs: 170
  • Evaluation Q&A pairs: 0

Hyperparameters

  • Max length: 2048
  • Epochs: 3.0
  • Learning rate: 3e-05
  • Batch size: 8
  • Grad accum: 2
  • Seed: 42
  • System prompt: "You are a helpful assistant. Answer the user's question accurately and concisely."
  • Weight decay: 0.1
  • Adam beta1: 0.9
  • Adam beta2: 0.95
  • Adam epsilon: 1e-08
  • Warmup steps: 0
  • Max grad norm: 1.0
  • Bf16: true
  • Tf32: true
  • Gradient checkpointing: true
  • Learning rate schedule: "cosine_with_0.1_minimum"
  • Optimizer: "adamw_torch"
  • Logging steps: 1
  • Save strategy: "epoch"
  • Save total limit: 3
  • Data seed: 42
  • Report to: "none"
  • Remove unused columns: false

Evaluation

  • dates_recall: 1.0
  • deterministic_pass_rate: 0.0
  • list_recall: 0.04358974358974359
  • numbers_recall: 0.7576923076923077
  • refusal_rate: 0.0
  • total: 26
  • unsupported_value_rate: 0.4230769230769231
  • urls_recall: 1.0

The external evaluation suite qualifies server-side behavior only. The exact converted artifact still requires browser and target-device validation.

MLC availability

MLC artifacts using q4f16_1 quantization are available under mlc/.

Limitations

The checkpoint can produce incorrect, incomplete, stale, or memorized content. The private training corpus is website-specific, and evaluation results do not establish general capability or production safety.

Provenance artifacts

  • research/run-manifest.json
  • research/training-config.json
  • research/publication-inventory.json
  • research/conversion-manifest.json
  • research/evaluation/metrics.json
  • research/evaluation/evaluation-manifest.json
Downloads last month
58
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kanha-AI/kanha-kanha.ai-1.7b-pit-v1

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1264)
this model