Safetensors
mpt
Krutrim
language-model
custom_code
krutrim-admin commited on
Commit
58c49dc
·
verified ·
1 Parent(s): 8c8ed60

Created Readme

Browse files
Files changed (1) hide show
  1. README.md +172 -0
README.md ADDED
@@ -0,0 +1,172 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ - hi
5
+ - bn
6
+ - mr
7
+ - te
8
+ - ta
9
+ - kn
10
+ - ml
11
+ - gu
12
+ - as
13
+ - pa
14
+ license: unknown
15
+ tags:
16
+ - Krutrim
17
+ - language-model
18
+ ---
19
+ # Krutrim-1
20
+
21
+ ## Model Overview
22
+ Krutrim Large Language Model (LLM) is a 2 trillion token multilingual foundation model designed to serve Indian demographic needs through equitable representation of the country's array of native tongues. Training data incorporates the largest known Indic language dataset, mitigating associated data scarcity obstacles that encumber model parity across dialects. Evaluations demonstrate Krutrim's strong performance on Indic language benchmarks, surpassing or at par with state-of-the-art models despite being significantly smaller in training flops. Krutrim LLM also matches or exceeds standards set on English benchmarks by models trained on comparable flops (e.g. vs LLAMA-2 on 10 out of 16 tasks with average score of 0.57 vs 0.55 of LLAMA-2), evidencing flexible multilingual fluency. Through intentional design choices that redress endemic data imbalances, Krutrim LLM signifies meaningful progress in the pursuit of ethical, globally representative AI foundation models.
23
+
24
+ ## Key Features
25
+ - 7B parameter dense transformer model comparable similarly sized LLama-2 model;
26
+ - Natively multilingual delivering best-in-class performance for a 7B mdoel on Indic benchmarks;
27
+ - Exceeds performance of similar sized models on multilingual Indic generation tasks including creative writing, summarization, and translation;
28
+ - Available in both pre-trained and instruction-tuned versions
29
+
30
+ ## Model Developer
31
+ - OLA Krutrim Team
32
+
33
+ ## Model Dates
34
+ - Krutrim-1 was trained between Oct 2023 and Nov 2023.
35
+
36
+ ## Release History
37
+
38
+ | Model Name | Release Date |Release Note | Reference|
39
+ |------------|-------------|-------------|-------------|
40
+ | Krutrim-1-Base | 2024-01-31 | Trained from scratch | [Here](https://huggingface.co/krutrim-ai-labs/Krutrim-1-base)
41
+ | Krutrim-1-Instruct | 2024-01-31 | SFT on Krutrim-1-Base |[Here](https://huggingface.co/krutrim-ai-labs/Krutrim-1-instruct)
42
+
43
+
44
+ ## Data Freshness
45
+ - The dataset includes information up to April 2023.
46
+
47
+ ## Model Architecture
48
+ - Layers: 32
49
+ - Max Sequence Length: 4096
50
+ - Hidden Dimension: 4608
51
+ - Head Dimension: 96
52
+ - Number of Heads: 48
53
+ - Number of KV-Heads: 8 (GQA)
54
+ - Vocabulary Size: 70400
55
+ - Architecture Type: Transformer Decoder (Auto-regressive Language Model)
56
+
57
+ ## Evaluation Results
58
+
59
+ ### English Comparison between Krutrim-1 and Llama2Chat (Benchmarks run on `llm_foundry`)
60
+
61
+ | Task | Llama2Chat | Krutrim-1-7B |
62
+ |--------------------|--------------|------------|
63
+ | arc | 0.517 | **0.557** |
64
+ | bigbench | **0.359** | 0.330 |
65
+ | boolq | **0.803** | 0.843 |
66
+ | copa | 0.78 | **0.82** |
67
+ | hellaswag | **0.754** | 0.740 |
68
+ | jeopardy | 0.306 | **0.286** |
69
+ | lambadaopenai | **0.695** | 0.682 |
70
+ | logiqa | 0.332 | **0.3195** |
71
+ | mathqa | **0.436** | 0.440 |
72
+ | mmlu | 0.472 | **0.495** |
73
+ | openbookqa | 0.44 | **0.464** |
74
+ | piqa | **0.7601** | 0.7726 |
75
+ | simplearithmetic | 0.160 | **0.077** |
76
+ | squad | 0.3565 | **0.369** |
77
+ | winograd | **0.8645** | 0.828 |
78
+ | winogrande | 0.681 | **0.697** |
79
+ | **average** | **0.54** | **0.54** |
80
+
81
+
82
+ ### Benchmarks
83
+
84
+ | Model | bn | gu | hi | kn | ml | mr | ta | te |
85
+ |------------------|------|------|------|------|------|------|------|------|
86
+ | **IndicCOPA** | | | | | | | | |
87
+ | Krutrim-1-7B | 0.89 | 0.83 | 0.86 | 0.88 | 0.88 | 0.87 | 0.89 | 0.89 |
88
+ | GPT-3.5 | 0.77 | 0.73 | 0.77 | 0.74 | 0.75 | 0.70 | 0.72 | 0.75 |
89
+ | Airawata | - | - | 0.74 | - | - | - | - | - |
90
+ | Kan-LLaMA | - | - | - | 0.74 | - | - | - | - |
91
+ | Tam-LLaMA | - | - | - | - | - | - | 0.77 | - |
92
+ | **IndicQA** | | | | | | | | |
93
+ | Krutrim-1-7B | 0.65 | 0.64 | 0.64 | 0.60 | 0.66 | 0.58 | 0.75 | 0.83 |
94
+ | Airawata | - | - | 0.62 | - | - | - | - | - |
95
+ | Kan-LLaMA | - | - | - | 0.52 | - | - | - | - |
96
+ | Tam-LLaMA | - | - | - | - | - | - | 0.35 | - |
97
+ | **IndicSentiment**| | | | | | | | |
98
+ | Krutrim-1-7B | 0.95 | 0.96 | 0.96 | 0.95 | 0.96 | 0.97 | 0.94 | 0.95 |
99
+ | GPT-3.5 | 0.50 | 0.81 | 0.96 | 0.60 | 0.75 | 0.88 | 0.51 | 0.53 |
100
+ | Airawata | - | - | 0.84 | - | - | - | - | - |
101
+ | Kan-LLaMA | - | - | - | 0.85 | - | - | - | - |
102
+ | Tam-LLaMA | - | - | - | - | - | - | 0.78 | - |
103
+ | **IndicTranslation**| | | | | | | | |
104
+ | Krutrim-1-7B | 0.88 | 0.89 | 0.95 | 0.88 | 0.89 | 0.92 | - | 0.88 |
105
+ | Airawata | - | - | 0.94 | - | - | - | - | - |
106
+ | Kan-LLaMA | - | - | - | 0.59 | - | - | - | - |
107
+ | **IndicXParaphrase**| | | | | | | | |
108
+ | Krutrim-1-7B | 0.91 | - | 0.97 | 0.82 | 0.90 | 0.94 | - | 0.61 |
109
+ | Airawata | - | - | 0.60 | - | - | - | - | - |
110
+ | Kan-LLaMA | - | - | - | 0.59 | - | - | - | - |
111
+
112
+ ## Usage
113
+
114
+ To run this model, do this:
115
+ ```
116
+ git clone https://github.com/ola-krutrim/Krutrim-1-7B.git
117
+ cd Krutrim-1-7B
118
+ pip install -r requirements.txt
119
+ ```
120
+
121
+ To test the base model, you can run
122
+ ```
123
+ python inference/inference.py
124
+ ```
125
+
126
+ To test batch inference of instruct model, you can run
127
+ ```
128
+ python inference/batch_inference.py
129
+ ```
130
+
131
+ To use the instruct model, you can load it with `AutoModelForCausalLM` as follows:
132
+ ```
133
+ import torch
134
+ from transformers import AutoModelForCausalLM, AutoTokenizer
135
+
136
+ model_id = "krutrim-ai-labs/Krutrim-1-base"
137
+ # Load model and tokenizer
138
+ model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16, trust_remote_code=True)
139
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
140
+
141
+ prompt = "Hello"
142
+
143
+ inputs = tokenizer(prompt, return_tensors='pt')
144
+ inputs.pop("token_type_ids", None)
145
+
146
+ # Generate response
147
+ outputs = model.generate(
148
+ **inputs,
149
+ max_length=5
150
+ )
151
+
152
+ response = tokenizer.decode(outputs[0])
153
+ print(response)
154
+ ```
155
+
156
+ ## Limitations
157
+ The model was trained on a dataset that includes content from the internet, which may contain toxic language, biases, and unsafe content. As a result, the model may:
158
+ - Amplify biases present in the training data
159
+ - Generate toxic responses, especially when prompted with toxic inputs
160
+ - Provide inaccurate, incomplete, or redundant answers
161
+ - Generate responses in languages inconsistent with the prompt
162
+
163
+ ## License
164
+ TBD
165
+
166
+ ## Ethical Considerations
167
+ - The model may produce biased or offensive outputs based on its training data.
168
+ - Users should apply human oversight when using the model for decision-making in sensitive areas.
169
+ - While safeguards have been implemented, the model may still generate socially undesirable text in certain contexts.
170
+
171
+ ## Contact
172
+ TBD