Bhishaj commited on
Commit
2bd7cae
·
1 Parent(s): 43240ae

🐛 Fix: Restore Space YAML configuration

Browse files
Files changed (1) hide show
  1. README.md +11 -64
README.md CHANGED
@@ -1,66 +1,13 @@
1
  ---
2
- language:
3
- - en
4
- - hi
5
- license: llama3.2
6
- tags:
7
- - legal
8
- - unsloth
9
- - turboquant
10
- - gguf
11
- - edge-ai
12
- datasets:
13
- - Techmaestro369/indian-legal-texts-finetuning
14
- - bharatgenai/BhashaBench-Legal
15
  ---
16
-
17
- # ⚖️ Vidhik AI: Sovereign Legal SLM (1B)
18
-
19
- ## Model Summary
20
- Vidhik AI is a highly optimized, domain-specific Small Language Model (SLM) engineered for the Indian Judiciary and MSME sector. Fine-tuned on a 1B parameter base, it specializes in drafting formal legal notices (e.g., MSMED Act delayed payments) and navigating complex Indian officialese.
21
-
22
- **Developer:** Bhishaj Technologies (Gaurav)
23
- **Base Model:** Llama-3.2-1B-Instruct
24
- **Quantization:** 4-bit GGUF (Q4_K_M)
25
-
26
- ## 🛠️ Training & MLOps Architecture
27
- To bypass local hardware constraints, the model was trained using a hybrid cloud-edge pipeline:
28
- * **Compute:** Kaggle Dual T4 GPUs (32GB VRAM)
29
- * **Optimization:** Unsloth for 70% VRAM reduction during fine-tuning.
30
- * **Method:** PEFT/QLoRA instruction fine-tuning on `indian-legal-texts-finetuning`.
31
- * **Guardrails:** Model is trained with strict negative stop-sequences and deterministic decoding (`Temperature = 0.0`) to prevent MCQ-loop hallucinations.
32
-
33
- ## ⚡ Edge Deployment & Google TurboQuant
34
- This model is specifically compiled to run on legacy/constrained hardware (e.g., NVIDIA GTX 1050 4GB).
35
-
36
- By utilizing **Google TurboQuant**, the model compresses the KV-cache to 3-bits during runtime, allowing for 128k context windows (essential for long Indian government gazettes) without triggering OOM (Out of Memory) crashes, maintaining a throughput of ~24.5 tokens/sec.
37
-
38
- ### Python Usage (TurboQuant Enabled)
39
- ```python
40
- import torch
41
- from transformers import AutoModelForCausalLM, AutoTokenizer
42
- from turboquant import TurboQuantCache
43
-
44
- repo_id = "Bhishaj/Vidhik-Llama-1B-GGU"
45
- tokenizer = AutoTokenizer.from_pretrained(repo_id)
46
- model = AutoModelForCausalLM.from_pretrained(repo_id, device_map="cuda")
47
-
48
- # Initialize TurboQuant 4-bit Cache for 4GB VRAM support
49
- tq_cache = TurboQuantCache(bits=4, compute_device="cuda")
50
-
51
- prompt = "TASK: Draft a formal legal notice for my client 'M/s Vidhik Electronics' under MSMED Act Sections 15 & 16."
52
- inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
53
-
54
- with torch.no_grad():
55
- outputs = model.generate(
56
- **inputs,
57
- past_key_values=tq_cache,
58
- max_new_tokens=512,
59
- temperature=0.0
60
- )
61
-
62
- print(tokenizer.decode(outputs[0], skip_special_tokens=True))
63
- ```
64
-
65
- ## 📊 Evaluation
66
- Evaluated against **BhashaBench-Legal (BBL)** to ensure alignment with Indian judicial service standards and formal legal tonality.
 
1
  ---
2
+ title: Vidhik AI Legal Assistant
3
+ emoji: ⚖️
4
+ colorFrom: blue
5
+ colorTo: indigo
6
+ sdk: gradio
7
+ sdk_version: 5.16.0
8
+ app_file: app.py
9
+ pinned: false
10
+ python_version: 3.11
 
 
 
 
11
  ---
12
+ Vidhik AI Legal Assistant
13
+ Sovereign Legal SLM running on Transformers-native GGUF.