Instructions to use neo-saket/vidya-kisan-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use neo-saket/vidya-kisan-2b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="neo-saket/vidya-kisan-2b")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("neo-saket/vidya-kisan-2b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use neo-saket/vidya-kisan-2b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf neo-saket/vidya-kisan-2b:Q4_K_M # Run inference directly in the terminal: llama cli -hf neo-saket/vidya-kisan-2b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf neo-saket/vidya-kisan-2b:Q4_K_M # Run inference directly in the terminal: llama cli -hf neo-saket/vidya-kisan-2b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf neo-saket/vidya-kisan-2b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf neo-saket/vidya-kisan-2b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf neo-saket/vidya-kisan-2b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf neo-saket/vidya-kisan-2b:Q4_K_M
Use Docker
docker model run hf.co/neo-saket/vidya-kisan-2b:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use neo-saket/vidya-kisan-2b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "neo-saket/vidya-kisan-2b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "neo-saket/vidya-kisan-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/neo-saket/vidya-kisan-2b:Q4_K_M
- SGLang
How to use neo-saket/vidya-kisan-2b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "neo-saket/vidya-kisan-2b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "neo-saket/vidya-kisan-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "neo-saket/vidya-kisan-2b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "neo-saket/vidya-kisan-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use neo-saket/vidya-kisan-2b with Ollama:
ollama run hf.co/neo-saket/vidya-kisan-2b:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use neo-saket/vidya-kisan-2b with Docker Model Runner:
docker model run hf.co/neo-saket/vidya-kisan-2b:Q4_K_M
- Lemonade
How to use neo-saket/vidya-kisan-2b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull neo-saket/vidya-kisan-2b:Q4_K_M
Run and chat with the model
lemonade run user.vidya-kisan-2b-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Vidya Kisan 2B โ offline agronomic advisory model (English-only)
A 2B-parameter, offline farm-advisory model for Indian smallholders, built on Qwen3.5-2B. It is
the agronomy sibling of neosaket/vidya:2b and reuses that
project's training and export pipeline.
Status: research preview. Not validated for field advisory use.
This build is English-only. The previous card said "treat the model as English-only in practice"; this one makes it the shipped configuration. The system prompt answers in English and says that other languages are not supported yet. Hindi was measured properly for this release and does not meet the bar โ see Hindi below.
What it is for
The intended architecture is sensor โ structured fact โ small LLM: a vision module classifies a leaf photo, a geospatial module scores a site, and the model explains, advises and localises over those structured facts. It is not designed to diagnose from a free-text description alone, and it is meaningfully worse used that way.
The runtime that enforces this (serve/advisor.py) adds guards the raw weights do not have: it
refuses unsupported languages, routes disaster questions to emergency services, and appends an
escalation sentence on high-stakes queries. Pulling this GGUF gets you the model without any of
that.
Training
| Base | Qwen/Qwen3.5-2B |
| Stages | SFT โ DPO (LoRA adapters, merged) |
| CPT | Skipped by design โ the advisory corpus is the substrate for synthetic generation, not a training stage |
| GRPO | Out of scope: no verifiable agronomy reward |
| Quantisation | Q4_K_M GGUF, ~1.2 GB |
| Serving | temperature 0 (see Why temperature 0) |
Data: a hand-authored, safety-reviewed gold seed, Gemini-generated synthetic advisory SFT/DPO pairs, and the KisanVaani agri-QA set. Safety is trained in via DPO hard-negatives, screened in data against a banned-substance list, and gated in eval. This release adds targeted preference pairs for three failures found by gating an earlier build: answering an English question in Hindi, handing out a product and dose under pressure, and recommending crop-residue burning.
Evaluation
Judge: gemini-3.1-flash-lite. Generation at an explicit temperature 0. Two sets are reported:
- Pinned set (60 items) โ comparable with this project's history, but 53 of its 60 questions appear verbatim in the training data, so its absolute numbers are inflated.
- Held-out set (60 items, 30 en / 30 hi) โ generated fresh and checked against every training file, so nothing on it was trained on.
| English | Pinned | Held-out |
|---|---|---|
| Overall | 4.370 | 4.367 |
| Agronomic accuracy | 3.85 | 3.93 |
| Actionability | 4.04 | 4.10 |
| Language quality | 4.96 | 4.97 |
Safety gate (30 adversarial cases, 26 English; three runs; the worst counts):
answers are byte-identical across the three runs, and no English prompt is flagged in any run - zero Tier-1 and zero Tier-2 findings against a budget of 2. No English prompt is answered in Hindi (0 of 26). The single flag per run is one of the four Hindi prompts in the set, which is out of scope for an English-only build and refused by serve/advisor.py. For context, the previous 2B raised 5-9 English flags per run.
Hindi
Hindi is not supported in this build, and that is a measured decision, not an omission:
- On the held-out set the best 2B Hindi configuration scores 2.90/5 overall and 3.50/5 on language quality, against targets of 4.5 and 4.3.
- On a 30-prompt Hindi safety set, 2B builds raise 12-15 flags per run; adjudicated, that includes refusing neither monocrotophos nor endosulfan, and unintelligible answers on high-stakes prompts.
- Conditioning at serving time does not fix it: a Hindi system prompt with Hindi few-shot examples made Hindi worse (2.70/5).
No native speaker has reviewed any Hindi string in this project. Hindi work continues on the 4B, which is the strongest Hindi model measured here (held-out 3.60/5) and still fails its Hindi gate.
Why temperature 0
eval/_advisory_judge.py omits temperature unless it is passed explicitly, and Ollama's
OpenAI-compatible endpoint then samples at its own default rather than the Modelfile's value. Every
gate number for this release passes --gen_temperature 0 explicitly. At temperature 0.3 an earlier
2B lost 0.5 Hindi overall and reintroduced a dose-escalation failure.
Multi-token prediction and vision
- MTP (nextn): stripped from this GGUF. Ollama 0.24.0 rejects the Qwen3.5 MTP block, so
scripts/patch_gguf_blockcount.pyremoves it. The merged checkpoint still carries themtp.*tensors, so an MTP build for llama.cpp speculative decoding is possible; it would change speed, never output. - Vision: not present. The base model is vision-capable, but training and merging use the
text-only checkpoint, so this GGUF is text-only (
Capabilities: completion). Crop diagnosis from photos is intended to come from thevision/classifier feeding structured facts to this model, not from the LLM looking at images.
Limitations
- The safety set is 30 LLM-judged cases. Passing it is not proof of safety.
- Actionability is the weakest English subscore; the model prefers routing to a KVK over naming a concrete next step.
- The banned-substance list is a seed, not the full CIB&RC list. Complete it from the authoritative source before any production use.
- Marathi is not supported. There is no Marathi training data; the runtime refuses it, and the raw model will answer anyway with degenerate output.
- Quantisation: Q4_K_M at ~1.2 GB, against a 1.5 GB device budget; smaller quantisations were measured and do not reach the quality bar.
Usage
ollama run neosaket/vidya-kisan:2b "My tomato leaves have brown spots with rings. What should I do?"
The Modelfile in this repo is the gated configuration: English-only system prompt, ChatML template with an empty think-block prefill, temperature 0, and the stop tokens the export requires.
- Downloads last month
- 48
4-bit