--- language: - en - hi license: llama3.2 tags: - legal - unsloth - turboquant - gguf - edge-ai datasets: - Techmaestro369/indian-legal-texts-finetuning - bharatgenai/BhashaBench-Legal --- # ⚖️ Vidhik AI: Sovereign Legal SLM (1B) ## Model Summary Vidhik AI is a highly optimized, domain-specific Small Language Model (SLM) engineered for the Indian Judiciary and MSME sector. Fine-tuned on a 1B parameter base, it specializes in drafting formal legal notices (e.g., MSMED Act delayed payments) and navigating complex Indian officialese. **Developer:** Bhishaj Technologies (Gaurav) **Base Model:** Llama-3.2-1B-Instruct **Quantization:** 4-bit GGUF (Q4_K_M) ## 🛠️ Training & MLOps Architecture To bypass local hardware constraints, the model was trained using a hybrid cloud-edge pipeline: * **Compute:** Kaggle Dual T4 GPUs (32GB VRAM) * **Optimization:** Unsloth for 70% VRAM reduction during fine-tuning. * **Method:** PEFT/QLoRA instruction fine-tuning on `indian-legal-texts-finetuning`. * **Guardrails:** Model is trained with strict negative stop-sequences and deterministic decoding (`Temperature = 0.0`) to prevent MCQ-loop hallucinations. ## ⚡ Edge Deployment & Google TurboQuant This model is specifically compiled to run on legacy/constrained hardware (e.g., NVIDIA GTX 1050 4GB). By utilizing **Google TurboQuant**, the model compresses the KV-cache to 3-bits during runtime, allowing for 128k context windows (essential for long Indian government gazettes) without triggering OOM (Out of Memory) crashes, maintaining a throughput of ~24.5 tokens/sec. ### Python Usage (TurboQuant Enabled) ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer from turboquant import TurboQuantCache repo_id = "Bhishaj/Vidhik-Llama-1B-GGU" tokenizer = AutoTokenizer.from_pretrained(repo_id) model = AutoModelForCausalLM.from_pretrained(repo_id, device_map="cuda") # Initialize TurboQuant 4-bit Cache for 4GB VRAM support tq_cache = TurboQuantCache(bits=4, compute_device="cuda") prompt = "TASK: Draft a formal legal notice for my client 'M/s Vidhik Electronics' under MSMED Act Sections 15 & 16." inputs = tokenizer(prompt, return_tensors="pt").to("cuda") with torch.no_grad(): outputs = model.generate( **inputs, past_key_values=tq_cache, max_new_tokens=512, temperature=0.0 ) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` ## 📊 Evaluation Evaluated against **BhashaBench-Legal (BBL)** to ensure alignment with Indian judicial service standards and formal legal tonality.