Vidya Kisan 2B โ€” offline agronomic advisory model (English-only)

A 2B-parameter, offline farm-advisory model for Indian smallholders, built on Qwen3.5-2B. It is the agronomy sibling of neosaket/vidya:2b and reuses that project's training and export pipeline.

Status: research preview. Not validated for field advisory use.

This build is English-only. The previous card said "treat the model as English-only in practice"; this one makes it the shipped configuration. The system prompt answers in English and says that other languages are not supported yet. Hindi was measured properly for this release and does not meet the bar โ€” see Hindi below.

What it is for

The intended architecture is sensor โ†’ structured fact โ†’ small LLM: a vision module classifies a leaf photo, a geospatial module scores a site, and the model explains, advises and localises over those structured facts. It is not designed to diagnose from a free-text description alone, and it is meaningfully worse used that way.

The runtime that enforces this (serve/advisor.py) adds guards the raw weights do not have: it refuses unsupported languages, routes disaster questions to emergency services, and appends an escalation sentence on high-stakes queries. Pulling this GGUF gets you the model without any of that.

Training

Base Qwen/Qwen3.5-2B
Stages SFT โ†’ DPO (LoRA adapters, merged)
CPT Skipped by design โ€” the advisory corpus is the substrate for synthetic generation, not a training stage
GRPO Out of scope: no verifiable agronomy reward
Quantisation Q4_K_M GGUF, ~1.2 GB
Serving temperature 0 (see Why temperature 0)

Data: a hand-authored, safety-reviewed gold seed, Gemini-generated synthetic advisory SFT/DPO pairs, and the KisanVaani agri-QA set. Safety is trained in via DPO hard-negatives, screened in data against a banned-substance list, and gated in eval. This release adds targeted preference pairs for three failures found by gating an earlier build: answering an English question in Hindi, handing out a product and dose under pressure, and recommending crop-residue burning.

Evaluation

Judge: gemini-3.1-flash-lite. Generation at an explicit temperature 0. Two sets are reported:

  • Pinned set (60 items) โ€” comparable with this project's history, but 53 of its 60 questions appear verbatim in the training data, so its absolute numbers are inflated.
  • Held-out set (60 items, 30 en / 30 hi) โ€” generated fresh and checked against every training file, so nothing on it was trained on.
English Pinned Held-out
Overall 4.370 4.367
Agronomic accuracy 3.85 3.93
Actionability 4.04 4.10
Language quality 4.96 4.97

Safety gate (30 adversarial cases, 26 English; three runs; the worst counts): answers are byte-identical across the three runs, and no English prompt is flagged in any run - zero Tier-1 and zero Tier-2 findings against a budget of 2. No English prompt is answered in Hindi (0 of 26). The single flag per run is one of the four Hindi prompts in the set, which is out of scope for an English-only build and refused by serve/advisor.py. For context, the previous 2B raised 5-9 English flags per run.

Hindi

Hindi is not supported in this build, and that is a measured decision, not an omission:

  • On the held-out set the best 2B Hindi configuration scores 2.90/5 overall and 3.50/5 on language quality, against targets of 4.5 and 4.3.
  • On a 30-prompt Hindi safety set, 2B builds raise 12-15 flags per run; adjudicated, that includes refusing neither monocrotophos nor endosulfan, and unintelligible answers on high-stakes prompts.
  • Conditioning at serving time does not fix it: a Hindi system prompt with Hindi few-shot examples made Hindi worse (2.70/5).

No native speaker has reviewed any Hindi string in this project. Hindi work continues on the 4B, which is the strongest Hindi model measured here (held-out 3.60/5) and still fails its Hindi gate.

Why temperature 0

eval/_advisory_judge.py omits temperature unless it is passed explicitly, and Ollama's OpenAI-compatible endpoint then samples at its own default rather than the Modelfile's value. Every gate number for this release passes --gen_temperature 0 explicitly. At temperature 0.3 an earlier 2B lost 0.5 Hindi overall and reintroduced a dose-escalation failure.

Multi-token prediction and vision

  • MTP (nextn): stripped from this GGUF. Ollama 0.24.0 rejects the Qwen3.5 MTP block, so scripts/patch_gguf_blockcount.py removes it. The merged checkpoint still carries the mtp.* tensors, so an MTP build for llama.cpp speculative decoding is possible; it would change speed, never output.
  • Vision: not present. The base model is vision-capable, but training and merging use the text-only checkpoint, so this GGUF is text-only (Capabilities: completion). Crop diagnosis from photos is intended to come from the vision/ classifier feeding structured facts to this model, not from the LLM looking at images.

Limitations

  • The safety set is 30 LLM-judged cases. Passing it is not proof of safety.
  • Actionability is the weakest English subscore; the model prefers routing to a KVK over naming a concrete next step.
  • The banned-substance list is a seed, not the full CIB&RC list. Complete it from the authoritative source before any production use.
  • Marathi is not supported. There is no Marathi training data; the runtime refuses it, and the raw model will answer anyway with degenerate output.
  • Quantisation: Q4_K_M at ~1.2 GB, against a 1.5 GB device budget; smaller quantisations were measured and do not reach the quality bar.

Usage

ollama run neosaket/vidya-kisan:2b "My tomato leaves have brown spots with rings. What should I do?"

The Modelfile in this repo is the gated configuration: English-only system prompt, ChatML template with an empty think-block prefill, temperature 0, and the stop tokens the export requires.

Downloads last month
48
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for neo-saket/vidya-kisan-2b

Finetuned
Qwen/Qwen3.5-2B
Quantized
(232)
this model