robinhad commited on
Commit
78a9356
·
1 Parent(s): ffd47fe

Add files using upload-large-folder tool

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,210 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: gemma
3
+ library_name: transformers
4
+ pipeline_tag: image-text-to-text
5
+ base_model: lapa-llm/lapa-12b-pt
6
+ language:
7
+ - uk
8
+ - en
9
+ ---
10
+
11
+ # Gemma 3 model card
12
+
13
+ ## Model Information
14
+
15
+ Introducing Lapa LLM v0.1.2 — the most efficient Ukrainian open-source language model
16
+
17
+ ### Description
18
+
19
+ > Demo page: https://huggingface.co/spaces/lapa-llm/lapa
20
+ > Link to Lapa Models: https://huggingface.co/collections/lapa-llm/lapa-v012-release
21
+ > Quantized versions: https://huggingface.co/lapa-llm/lapa-v0.1.2-instruct-GGUF
22
+
23
+
24
+ Today, we proudly present Lapa LLM — a cutting-edge open large language model based on Gemma-3-12B with a focus on Ukrainian language processing. The project is the result of many months of work by a team of Ukrainian researchers in artificial intelligence from the Ukrainian Catholic University, AGH University of Krakow, Igor Sikorsky Kyiv Polytechnic Institute, and Lviv Polytechnic, who united to create the best model for Ukrainian language processing.
25
+
26
+ The model is named in honor of [Valentyn Lapa](https://de.wikipedia.org/wiki/Walentyn_Lapa), who together with [Oleksiy Ivakhnenko](https://uk.wikipedia.org/wiki/%D0%86%D0%B2%D0%B0%D1%85%D0%BD%D0%B5%D0%BD%D0%BA%D0%BE_%D0%9E%D0%BB%D0%B5%D0%BA%D1%81%D1%96%D0%B9_%D0%93%D1%80%D0%B8%D0%B3%D0%BE%D1%80%D0%BE%D0%B2%D0%B8%D1%87) created the Group Method of Data Handling, which is a predecessor to Deep Learning [(source)](https://people.idsia.ch/~juergen/DeepLearning2July2014.pdf).
27
+
28
+ The project's goal is to create the best model for Ukrainian language processing with open datasets for pretraining and instruction tuning.
29
+
30
+ ### Key Achievements
31
+
32
+ **Best tokenizer for the Ukrainian language**
33
+
34
+ Thanks to a SOTA method for tokenizer adaptation developed by [Mykola Haltiuk](https://www.linkedin.com/in/mykola-haltiuk/) as part of this project, it was possible to replace 80,000 tokens out of 250,000 with Ukrainian ones without loss of model quality, thus making Lapa LLM the fastest model for working with the Ukrainian language. Compared to the original Gemma 3, for working with Ukrainian, the model requires 1.5 times fewer tokens, thus performing three times fewer computations to achieve better results.
35
+
36
+ **Most efficient instruction-tuned model on the market**
37
+
38
+ Our instruction version of the model in some benchmark categories is only slightly behind the current leader — [MamayLM](https://huggingface.co/spaces/INSAIT-Institute/mamaylm-v1-blog). The team is actively working on new datasets to further improve benchmark scores, which we plan to surpass in the v1.0 model.
39
+
40
+ ### Benchmark Results
41
+
42
+ - Best English-to-Ukrainian translator with a result of 33 BLEU on FLORES and vice versa, which allows for natural and cost-effective translation of new NLP datasets into Ukrainian
43
+ - One of the best models for image processing in Ukrainian in its size class, as measured on the MMZNO benchmark
44
+ - One of the best models for Summarization and Q&A, which means excellent performance for RAG
45
+ - Tests on propaganda and disinformation questions show the effectiveness of the filtering approach at the pretraining stage and during instruction fine-tuning
46
+
47
+ Model measurements and comparisons will be published as part of the Ukrainian LLM Leaderboard project; subscribe to the Telegram channel for further news.
48
+
49
+ **Leader in pretraining results**
50
+
51
+ Lapa LLM demonstrates the best performance in pretraining benchmarks for Ukrainian language processing, which opens opportunities for use by other researchers to adapt for their own tasks.
52
+
53
+ The model was trained on data evaluated by various quality assessment models - evaluation of propaganda and disinformation presence, readability, grammar assessment, etc. In the final stages of training, the model was trained on high-quality materials provided for commercial use by the Open Data division of Harvard Library.
54
+
55
+ **Maximum openness and transparency**
56
+
57
+ Unlike most available models, Lapa LLM is a maximally open project:
58
+ - The model is available for commercial use
59
+ - Approximately 25 datasets for model training have been published
60
+ - Methods for filtering and processing data are disclosed, including for detecting disinformation and propaganda
61
+ - Open source code for the model
62
+ - Documentation of the training process is available
63
+
64
+ This openness allows for the development of the Ukrainian NLP community and helps businesses obtain a tool for the most efficient Ukrainian language processing in terms of both computation and results.
65
+
66
+ ### Application Possibilities
67
+
68
+ Lapa LLM opens wide possibilities for:
69
+ - Processing sensitive documents without transferring data to external servers
70
+ - Working with Ukrainian texts taking into account cultural and historical context without code-switching to Russian or other languages
71
+ - Building RAG systems and chatbots that write in proper Ukrainian
72
+ - Developing specialized solutions through the ability to fine-tune for specific tasks
73
+ - Machine translation with the best translation quality from English to Ukrainian and vice versa among all models, including API providers
74
+
75
+ ### Next Steps
76
+
77
+ - Complete development of the reasoning model
78
+ - We are collecting community feedback on the model's performance, so we look forward to receiving it on GitHub or HuggingFace!
79
+ - Collecting additional datasets for image processing in Ukrainian
80
+ - Collecting additional datasets for instruction following and programming
81
+
82
+ ### Acknowledgment to Sponsors
83
+
84
+ The creation of Lapa LLM was made possible thanks to the support of our partners and sponsors, primarily the startup **Comand.AI**, which provided computational resources for training the model. We also want to thank the company **ELEKS**, which supported this project through a grant dedicated to the memory of Oleksiy Skrypnyk, and the startup **HuggingFace**, which provided a free corporate subscription to the team for storing models and datasets.
85
+
86
+ ### Links:
87
+
88
+ Try the model: https://huggingface.co/spaces/lapa-llm/lapa
89
+ Code: https://github.com/lapa-llm/lapa-llm
90
+
91
+ Subscribe to the Telegram channel for further news about the project: https://t.me/pehade_blog
92
+
93
+ ### Team
94
+
95
+ ### Inputs and outputs
96
+
97
+ - **Input:**
98
+ - Text string, such as a question, a prompt, or a document to be summarized
99
+ - Images, normalized to 896 x 896 resolution and encoded to 256 tokens
100
+ each
101
+ - Total input context of 128K tokens
102
+
103
+ - **Output:**
104
+ - Generated text in response to the input, such as an answer to a
105
+ question, analysis of image content, or a summary of a document
106
+ - Total output context of 8192 tokens
107
+
108
+ ### Usage
109
+
110
+ Below, there are some code snippets on how to get quickly started with running the model. First, install the Transformers library. Gemma 3 is supported starting from transformers 4.50.0.
111
+
112
+ ```sh
113
+ $ pip install -U transformers
114
+ ```
115
+
116
+ Then, copy the snippet from the section that is relevant for your use case.
117
+
118
+ #### Running with the `pipeline` API
119
+
120
+ You can initialize the model and processor for inference with `pipeline` as follows.
121
+
122
+ ```python
123
+ from transformers import pipeline
124
+ import torch
125
+
126
+ pipe = pipeline(
127
+ "image-text-to-text",
128
+ model="lapa-llm/lapa-v0.1.2-instruct",
129
+ device="cuda",
130
+ torch_dtype=torch.bfloat16
131
+ )
132
+ ```
133
+
134
+ With instruction-tuned models, you need to use chat templates to process our inputs first. Then, you can pass it to the pipeline.
135
+
136
+ ```python
137
+ messages = [
138
+ {
139
+ "role": "system",
140
+ "content": [{"type": "text", "text": "You are a helpful assistant."}]
141
+ },
142
+ {
143
+ "role": "user",
144
+ "content": [
145
+ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
146
+ {"type": "text", "text": "What animal is on the candy?"}
147
+ ]
148
+ }
149
+ ]
150
+
151
+ output = pipe(text=messages, max_new_tokens=200)
152
+ print(output[0]["generated_text"][-1]["content"])
153
+ # Okay, let's take a look!
154
+ # Based on the image, the animal on the candy is a **turtle**.
155
+ # You can see the shell shape and the head and legs.
156
+ ```
157
+
158
+ #### Running the model on a single / multi GPU
159
+
160
+ ```python
161
+ # pip install accelerate
162
+
163
+ from transformers import AutoProcessor, Gemma3ForConditionalGeneration
164
+ from PIL import Image
165
+ import requests
166
+ import torch
167
+
168
+ model_id = "lapa-llm/lapa-v0.1.2-instruct"
169
+
170
+ model = Gemma3ForConditionalGeneration.from_pretrained(
171
+ model_id, device_map="auto"
172
+ ).eval()
173
+
174
+ processor = AutoProcessor.from_pretrained(model_id)
175
+
176
+ messages = [
177
+ {
178
+ "role": "system",
179
+ "content": [{"type": "text", "text": "You are a helpful assistant."}]
180
+ },
181
+ {
182
+ "role": "user",
183
+ "content": [
184
+ {"type": "image", "image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"},
185
+ {"type": "text", "text": "Опиши зображення"}
186
+ ]
187
+ }
188
+ ]
189
+
190
+ inputs = processor.apply_chat_template(
191
+ messages, add_generation_prompt=True, tokenize=True,
192
+ return_dict=True, return_tensors="pt"
193
+ ).to(model.device, dtype=torch.bfloat16)
194
+
195
+ input_len = inputs["input_ids"].shape[-1]
196
+
197
+ with torch.inference_mode():
198
+ generation = model.generate(**inputs, max_new_tokens=100, do_sample=False)
199
+ generation = generation[0][input_len:]
200
+
201
+ decoded = processor.decode(generation, skip_special_tokens=True)
202
+ print(decoded)
203
+
204
+ # **Overall Impression:** The image is a close-up shot of a vibrant garden scene,
205
+ # focusing on a cluster of pink cosmos flowers and a busy bumblebee.
206
+ # It has a slightly soft, natural feel, likely captured in daylight.
207
+ ```
208
+
209
+ ### Citation
210
+ TBD
chat_template.jinja ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {{ bos_token }}
2
+ {%- if messages[0]['role'] == 'system' -%}
3
+ {%- if messages[0]['content'] is string -%}
4
+ {%- set first_user_prefix = messages[0]['content'] + '
5
+
6
+ ' -%}
7
+ {%- else -%}
8
+ {%- set first_user_prefix = messages[0]['content'][0]['text'] + '
9
+
10
+ ' -%}
11
+ {%- endif -%}
12
+ {%- set loop_messages = messages[1:] -%}
13
+ {%- else -%}
14
+ {%- set first_user_prefix = "" -%}
15
+ {%- set loop_messages = messages -%}
16
+ {%- endif -%}
17
+ {%- for message in loop_messages -%}
18
+ {%- if (message['role'] == 'user') != (loop.index0 % 2 == 0) -%}
19
+ {{ raise_exception("Conversation roles must alternate user/assistant/user/assistant/...") }}
20
+ {%- endif -%}
21
+ {%- if (message['role'] == 'assistant') -%}
22
+ {%- set role = "model" -%}
23
+ {%- else -%}
24
+ {%- set role = message['role'] -%}
25
+ {%- endif -%}
26
+ {{ '<start_of_turn>' + role + '
27
+ ' + (first_user_prefix if loop.first else "") }}
28
+ {%- if message['content'] is string -%}
29
+ {{ message['content'] | trim }}
30
+ {%- elif message['content'] is iterable -%}
31
+ {%- for item in message['content'] -%}
32
+ {%- if item['type'] == 'image' -%}
33
+ {{ '<start_of_image>' }}
34
+ {%- elif item['type'] == 'text' -%}
35
+ {{ item['text'] | trim }}
36
+ {%- endif -%}
37
+ {%- endfor -%}
38
+ {%- else -%}
39
+ {{ raise_exception("Invalid content type") }}
40
+ {%- endif -%}
41
+ {{ '<end_of_turn>
42
+ ' }}
43
+ {%- endfor -%}
44
+ {%- if add_generation_prompt -%}
45
+ {{'<start_of_turn>model
46
+ '}}
47
+ {%- endif -%}
config.json ADDED
@@ -0,0 +1,119 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Gemma3ForConditionalGeneration"
4
+ ],
5
+ "boi_token_index": 255999,
6
+ "dtype": "bfloat16",
7
+ "eoi_token_index": 256000,
8
+ "eos_token_id": 1,
9
+ "image_token_index": 262144,
10
+ "initializer_range": 0.02,
11
+ "mm_tokens_per_image": 256,
12
+ "model_type": "gemma3",
13
+ "text_config": {
14
+ "_sliding_window_pattern": 6,
15
+ "attention_bias": false,
16
+ "attention_dropout": 0.0,
17
+ "attn_logit_softcapping": null,
18
+ "bos_token_id": 2,
19
+ "dtype": "bfloat16",
20
+ "eos_token_id": 1,
21
+ "final_logit_softcapping": null,
22
+ "head_dim": 256,
23
+ "hidden_activation": "gelu_pytorch_tanh",
24
+ "hidden_size": 3840,
25
+ "initializer_range": 0.02,
26
+ "intermediate_size": 15360,
27
+ "layer_types": [
28
+ "sliding_attention",
29
+ "sliding_attention",
30
+ "sliding_attention",
31
+ "sliding_attention",
32
+ "sliding_attention",
33
+ "full_attention",
34
+ "sliding_attention",
35
+ "sliding_attention",
36
+ "sliding_attention",
37
+ "sliding_attention",
38
+ "sliding_attention",
39
+ "full_attention",
40
+ "sliding_attention",
41
+ "sliding_attention",
42
+ "sliding_attention",
43
+ "sliding_attention",
44
+ "sliding_attention",
45
+ "full_attention",
46
+ "sliding_attention",
47
+ "sliding_attention",
48
+ "sliding_attention",
49
+ "sliding_attention",
50
+ "sliding_attention",
51
+ "full_attention",
52
+ "sliding_attention",
53
+ "sliding_attention",
54
+ "sliding_attention",
55
+ "sliding_attention",
56
+ "sliding_attention",
57
+ "full_attention",
58
+ "sliding_attention",
59
+ "sliding_attention",
60
+ "sliding_attention",
61
+ "sliding_attention",
62
+ "sliding_attention",
63
+ "full_attention",
64
+ "sliding_attention",
65
+ "sliding_attention",
66
+ "sliding_attention",
67
+ "sliding_attention",
68
+ "sliding_attention",
69
+ "full_attention",
70
+ "sliding_attention",
71
+ "sliding_attention",
72
+ "sliding_attention",
73
+ "sliding_attention",
74
+ "sliding_attention",
75
+ "full_attention"
76
+ ],
77
+ "max_position_embeddings": 131072,
78
+ "model_type": "gemma3_text",
79
+ "num_attention_heads": 16,
80
+ "num_hidden_layers": 48,
81
+ "num_key_value_heads": 8,
82
+ "pad_token_id": 0,
83
+ "query_pre_attn_scalar": 256,
84
+ "rms_norm_eps": 1e-06,
85
+ "rope_parameters": {
86
+ "full_attention": {
87
+ "factor": 8.0,
88
+ "rope_theta": 1000000.0,
89
+ "rope_type": "linear"
90
+ },
91
+ "sliding_attention": {
92
+ "rope_theta": 10000.0,
93
+ "rope_type": "default"
94
+ }
95
+ },
96
+ "sliding_window": 1024,
97
+ "tie_word_embeddings": true,
98
+ "use_bidirectional_attention": false,
99
+ "use_cache": false,
100
+ "vocab_size": 262208
101
+ },
102
+ "tie_word_embeddings": true,
103
+ "transformers_version": "5.10.2",
104
+ "vision_config": {
105
+ "attention_dropout": 0.0,
106
+ "dtype": "bfloat16",
107
+ "hidden_act": "gelu_pytorch_tanh",
108
+ "hidden_size": 1152,
109
+ "image_size": 896,
110
+ "intermediate_size": 4304,
111
+ "layer_norm_eps": 1e-06,
112
+ "model_type": "siglip_vision_model",
113
+ "num_attention_heads": 16,
114
+ "num_channels": 3,
115
+ "num_hidden_layers": 27,
116
+ "patch_size": 14,
117
+ "vision_use_head": false
118
+ }
119
+ }
generation_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 2,
3
+ "cache_implementation": "hybrid",
4
+ "do_sample": true,
5
+ "eos_token_id": [
6
+ 1,
7
+ 106
8
+ ],
9
+ "pad_token_id": 0,
10
+ "top_k": 64,
11
+ "top_p": 0.95,
12
+ "transformers_version": "5.10.2"
13
+ }
merge_info.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "base_model": "lapa-llm/lapa-v0.1.2-instruct",
3
+ "adapter": "outputs/multilingual_crimea_dpo_lora_v2/adapter",
4
+ "merged_output": "outputs/multilingual_crimea_dpo_lora_v2_merged",
5
+ "torch_dtype": "bfloat16",
6
+ "safe_serialization": true
7
+ }
model-00001-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3dbf2bff4f6d2d8b25e3f0cebca3e10f78fc284538c91b416b9bb73527ed2050
3
+ size 3924929072
model-00002-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:50c2fcaf066c7d77c1e51e2f052e0abbb7c86f13707100830141e4b16b26b673
3
+ size 3916732256
model-00003-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6936a57f47e5fbae4191d393f2c5ca3b4ec061262446783bd17555251bae1a5d
3
+ size 3987503280
model-00004-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:10234fc6a4f5bf505314e53bb1d1553a1dce629d826ee16d20fedbd5457edc57
3
+ size 3987510440
model-00005-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fad4cdb282967c42df40da7863c1c9a7f38f2862992eeb9cd4e93d2c31a001c1
3
+ size 3916708232
model-00006-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c96c1e2a90784510a755e552883e3dd7128149f735c45000cf1f66c24ef0424e
3
+ size 3998644864
model-00007-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:04335090d6f15189ce63cc788052aff4b88ff5ff07ba465ba4c3ce8b45578d9f
3
+ size 642764096
model.safetensors.index.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:96c223ece326885174a741fcb28446f49757a4f491b1187cdeb312c13b98b1ba
3
+ size 37379975
tokenizer_config.json ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "boi_token": "<start_of_image>",
4
+ "bos_token": "<bos>",
5
+ "clean_up_tokenization_spaces": false,
6
+ "eoi_token": "<end_of_image>",
7
+ "eos_token": "<eos>",
8
+ "image_token": "<image_soft_token>",
9
+ "is_local": true,
10
+ "local_files_only": false,
11
+ "mask_token": "<mask>",
12
+ "model_max_length": 1000000000000000019884624838656,
13
+ "model_specific_special_tokens": {
14
+ "boi_token": "<start_of_image>",
15
+ "eoi_token": "<end_of_image>",
16
+ "image_token": "<image_soft_token>"
17
+ },
18
+ "pad_token": "<pad>",
19
+ "processor_class": "Gemma3Processor",
20
+ "sp_model_kwargs": null,
21
+ "spaces_between_special_tokens": false,
22
+ "tokenizer_class": "GemmaTokenizer",
23
+ "unk_token": "<unk>",
24
+ "use_default_system_prompt": false
25
+ }