OpceanAI commited on
Commit
262104e
·
verified ·
1 Parent(s): 6753a8a

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +73 -317
README.md CHANGED
@@ -3,335 +3,91 @@ license: apache-2.0
3
  language:
4
  - es
5
  - en
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6
  ---
7
 
8
- PASITA v1
9
 
10
- PASITA is a compact generative language model specialized in plain-text to Markdown (GFM) transformation with an emphasis on preserving the information contained in the input.
11
 
12
- It is a decoder-only Transformer based on the Llama architecture, trained entirely from scratch without a pretrained base model. PASITA is designed as a specialized text transformation model rather than a general-purpose conversational or instruction-following language model.
13
 
14
- Overview
15
 
16
- PASITA targets the following transformation:
 
 
 
 
 
 
 
 
 
 
17
 
18
- Plain text
19
- |
20
- v
21
- PASITA
22
- |
23
- v
24
- Markdown (GFM)
25
 
26
- The primary objective is to introduce Markdown structure while minimizing changes to the source content.
 
 
 
27
 
28
- PASITA is intended for use cases where preserving the original information is more important than aggressively formatting the input.
29
 
30
- Model Specifications
 
 
 
31
 
32
- Property| Value
33
- Model| PASITA v1
34
- Task| Plain text → Markdown (GFM)
35
- Architecture| LlamaForCausalLM
36
- Parameters| 88,099,584 (~88M)
37
- Hidden size| 768
38
- Layers| 12
39
- Intermediate size| 2048
40
- Attention heads| 12
41
- Key/value heads| 4
42
- Attention| Grouped Query Attention (GQA)
43
- Head dimension| 64
44
- Activation| SiLU / SwiGLU
45
- Normalization| RMSNorm
46
- RMSNorm epsilon| 1e-5
47
- Positional encoding| RoPE
48
- RoPE theta| 100,000
49
- Maximum context| 2,048 tokens
50
- Vocabulary| 16,384 tokens
51
- Word embeddings| Tied
52
- Attention bias| Disabled
53
- MLP bias| Disabled
54
- Parameter dtype| bfloat16
55
- Weight format| safetensors
56
- Number of tensors| 110
57
 
58
- Architecture
 
 
 
 
 
 
 
 
59
 
60
- PASITA uses a 12-layer decoder-only Transformer with a hidden dimension of 768 and an intermediate dimension of 2048.
61
 
62
- The attention mechanism uses Grouped Query Attention with 12 query heads and 4 key/value heads. This reduces the number of key/value projections and associated cache requirements while retaining multiple query heads.
63
 
64
- The model uses:
65
-
66
- - RMSNorm
67
- - SiLU activation
68
- - SwiGLU feed-forward blocks
69
- - Rotary Position Embeddings (RoPE)
70
- - A maximum sequence length of 2,048 tokens
71
- - Tied input and output embeddings
72
- - No attention or MLP biases
73
-
74
- The model is implemented using "LlamaForCausalLM".
75
-
76
- Tokenizer
77
-
78
- PASITA uses a custom Byte-Level BPE tokenizer with a vocabulary of 16,384 pieces.
79
-
80
- The tokenizer includes:
81
-
82
- <pad>
83
- <s>
84
- </s>
85
- <unk>
86
- <think>
87
- </think>
88
-
89
- The ByteLevel pre-tokenizer and decoder have been verified using exact round-trip tests.
90
-
91
- Measured tokenization efficiency is approximately:
92
-
93
- - 4.0 characters per token for Spanish and English
94
- - 2.0 characters per token for code
95
-
96
- Training Data
97
-
98
- PASITA was trained on approximately 80 million BPE tokens.
99
-
100
- The underlying corpus consists of human-produced Markdown obtained from public sources, including:
101
-
102
- - GoodWiki English
103
- - Spanish Wikipedia dumps
104
- - StackExchange Q&A
105
- - WikiHow
106
- - Cosmopedia
107
-
108
- The Markdown documents were deterministically degraded to produce noisy plain-text inputs.
109
-
110
- The resulting training data contains approximately 37% Spanish and 63% English.
111
-
112
- Dataset Splits
113
-
114
- Split| Tokens / prompts| Description
115
- SFT| 58M / 47k| Supervised fine-tuning
116
- DPO| 16M / 9k pairs| Preference optimization
117
- GRPO| 6M / 10.4k prompts| Group-relative optimization
118
- Benchmark| 2.5M / 2k| Held-out evaluation
119
-
120
- The benchmark split is disjoint from the training data by hash.
121
-
122
- Dataset quality controls included Markdown validity checks, hash and ROUGE-based deduplication, context-length validation, and multi-judge verification.
123
-
124
- Training
125
-
126
- PASITA v1 was trained using bfloat16 on NVIDIA B300 and A100 hardware.
127
-
128
- The training pipeline consisted of three stages.
129
-
130
- 1. Supervised Fine-Tuning
131
-
132
- SFT was performed for two epochs.
133
-
134
- Reported final metrics:
135
-
136
- - Loss: approximately 0.09
137
- - Token accuracy: 98.4%
138
-
139
- 2. Direct Preference Optimization
140
-
141
- DPO was performed for one epoch with:
142
-
143
- β = 0.1
144
-
145
- The reported final preference margin was approximately 4.2.
146
-
147
- The original DPO run contained partially spurious signals in some rejected examples. A subsequent DPO-v2 dataset was developed with more minimal edits and harder negatives.
148
-
149
- 3. Group Relative Policy Optimization
150
-
151
- GRPO was performed for 550 steps using groups of four generations.
152
-
153
- The reward design focused primarily on structural correctness and content fidelity.
154
-
155
- During the original training process, the fidelity component of the reward exhibited noise related to the decoder configuration, while structural rewards were learned more consistently.
156
-
157
- Evaluation
158
-
159
- Evaluation was performed greedily on held-out data.
160
-
161
- Bench-500
162
-
163
- Metric| Score
164
- GFM validity| 0.958
165
- Faithfulness| 0.930
166
- Semantic faithfulness| 0.880
167
- Table correctness| 1.000
168
-
169
- Bench-1000
170
-
171
- Metric| Score
172
- GFM validity| 0.956
173
- Faithfulness| 0.927
174
- Semantic faithfulness| 0.888
175
- Table correctness| 1.000
176
-
177
- Domain Results
178
-
179
- Domain| Score
180
- Code| 0.965
181
- OCR / transcripts| 0.953
182
- Documents| 0.946
183
- Mathematics| 0.981
184
- Tables| 0.945
185
- HTML| 1.000
186
- Control| 0.190
187
-
188
- The control category is the principal known weakness of the current version.
189
-
190
- Inference
191
-
192
- PASITA expects a prompt beginning with "CONVIERTE A MARKDOWN", followed by the source text and the "### Markdown:" separator.
193
-
194
- Example:
195
-
196
- prompt = (
197
- "CONVIERTE A MARKDOWN:\n"
198
- + text
199
- + "\n\n### Markdown:\n"
200
- )
201
-
202
- A recommended deterministic generation configuration is:
203
-
204
- generate(
205
- do_sample=False,
206
- max_new_tokens=int(input_length * 1.3 + 128)
207
- )
208
-
209
- The model is intended for deterministic transformation rather than conversational generation.
210
-
211
- Model Size and Runtime
212
-
213
- The model contains approximately 88 million parameters.
214
-
215
- The stored model weights occupy approximately:
216
-
217
- 176,211,400 bytes
218
-
219
- or approximately 168 MB.
220
-
221
- The weights are stored natively in the safetensors format.
222
-
223
- The model uses bfloat16 parameters, while CPU inference can be performed using a float32-loaded model when required by the runtime configuration.
224
-
225
- Approximate memory usage for the weights is around 180 MB, excluding runtime allocations and KV-cache requirements.
226
-
227
- The model is therefore small enough to be deployed on CPU systems and on GPUs with at least 2 GB of memory, subject to runtime overhead and the selected inference configuration.
228
-
229
- Weight Integrity
230
-
231
- The primary model file is:
232
-
233
- final_model/model.safetensors
234
-
235
- SHA-256 prefix:
236
-
237
- 664665e7ee2b0002
238
-
239
- Full hashes should be obtained from the corresponding release artifacts when verifying a downloaded model.
240
-
241
- Intended Use
242
-
243
- PASITA is designed for applications that require deterministic or near-deterministic conversion of unstructured text into Markdown.
244
-
245
- Potential applications include:
246
-
247
- - Plain-text document formatting
248
- - Markdown normalization pipelines
249
- - Documentation preprocessing
250
- - Text-to-Markdown conversion
251
- - Local document transformation
252
- - CPU-based Markdown processing
253
-
254
- PASITA is not designed or evaluated as a general-purpose conversational assistant.
255
-
256
- Known Limitations
257
-
258
- PASITA v1 has several known limitations.
259
-
260
- Content Preservation
261
-
262
- Although faithfulness is substantially stronger than its Markdown validity score alone would indicate, the model can still:
263
-
264
- - Truncate numerical values
265
- - Omit secondary information
266
- - Echo headings unnecessarily
267
- - Continue generating after completing the transformation
268
-
269
- For example, numerical content such as "23.3" may occasionally be reduced to "23".
270
-
271
- Short Inputs
272
-
273
- The current model is less reliable on very short inputs, particularly inputs consisting of only one to three lines.
274
-
275
- Strict Formatting Instructions
276
-
277
- The control benchmark remains substantially weaker than the other evaluated domains. PASITA should therefore not be assumed to reliably satisfy arbitrary, highly specific formatting constraints.
278
-
279
- Generation Termination
280
-
281
- Under greedy generation, PASITA can occasionally continue generating after the correct Markdown transformation has already been produced.
282
-
283
- Applications should therefore use an appropriate generation limit and explicitly handle end-of-sequence behavior.
284
-
285
- License
286
-
287
- PASITA v1 is released under the Apache License 2.0.
288
-
289
- Copyright © 2026 OpceanAI.
290
-
291
- The Apache License 2.0 permits use, reproduction, modification, distribution, and creation of derivative works subject to the terms and conditions of the license.
292
-
293
- The complete license text is provided in the ""LICENSE"" (LICENSE) file.
294
-
295
- The Apache License 2.0 applies to the PASITA project and its released model artifacts. Third-party datasets and source material used during training may be subject to separate licenses, attribution requirements, or other terms.
296
-
297
- Users intending to redistribute or commercially deploy systems trained on the underlying datasets should independently review the applicable licenses and attribution requirements, including the requirements associated with Wikipedia-derived material.
298
-
299
- Data Provenance
300
-
301
- The model weights were produced from a random initialization and were not initialized from a pretrained language model.
302
-
303
- Training data was derived from publicly available sources, including GoodWiki, Wikipedia, StackExchange, WikiHow, and Cosmopedia.
304
-
305
- The licensing status of the resulting model does not supersede the licensing requirements applicable to third-party source material.
306
-
307
- Project Status
308
-
309
- PASITA v1 should be considered an experimental specialized language model.
310
-
311
- The current results demonstrate that a compact model trained specifically for plain-text-to-Markdown transformation can achieve high content-fidelity scores on held-out evaluation data while remaining small enough for practical local inference.
312
-
313
- Future development is focused on improving:
314
-
315
- - Termination behavior
316
- - Control and instruction-following performance
317
- - Content preservation
318
- - Robustness on short inputs
319
- - Markdown structural correctness
320
- - Preference and reinforcement-learning data quality
321
-
322
- Summary
323
-
324
- PASITA v1 is an approximately 88M-parameter decoder-only Transformer trained from scratch specifically for plain-text-to-Markdown transformation.
325
-
326
- Its current Bench-1000 results are:
327
-
328
- GFM validity 95.6%
329
- Faithfulness 92.7%
330
- Semantic faith. 88.8%
331
- Table correctness 100.0%
332
-
333
- The model's primary strength is its ability to perform structured Markdown transformation while preserving most of the source information.
334
-
335
- Its primary weaknesses are strict control tasks, very short inputs, and occasional post-completion generation.
336
-
337
- PASITA is therefore best understood as a small, specialized text-transformation model, rather than a general-purpose LLM.
 
3
  language:
4
  - es
5
  - en
6
+ library_name: transformers
7
+ pipeline_tag: text-generation
8
+ tags:
9
+ - markdown
10
+ - text-to-markdown
11
+ - faithful-generation
12
+ - llama
13
+ - from-scratch
14
+ - spanish
15
+ - english
16
+ - ocr-postprocessing
17
+ - rag
18
+ - document-understanding
19
+ - tiny-llm
20
+ model-index:
21
+ - name: PASITA
22
+ results:
23
+ - task:
24
+ type: text-generation
25
+ dataset:
26
+ name: PASITA-bench-1000 (held-out, private)
27
+ type: custom
28
+ metrics:
29
+ - name: gfm_validity
30
+ type: accuracy
31
+ value: 0.956
32
+ - name: faithfulness
33
+ type: faithfulness
34
+ value: 0.927
35
+ - name: semantic_faithfulness
36
+ type: faithfulness
37
+ value: 0.888
38
+ - name: table_fidelity
39
+ type: accuracy
40
+ value: 1.0
41
  ---
42
 
43
+ # PASITA v1 — plain text → Markdown (fiel)
44
 
45
+ Modelo de lenguaje **decoder-only entrenado 100% desde cero** (sin modelo base), especializado en una sola tarea: convertir **texto plano en Markdown válido preservando la información**.
46
 
47
+ > Comportamiento de compilador, no de chatbot: añade estructura, no inventa contenido.
48
 
49
+ ## Arquitectura (laboratorio)
50
 
51
+ | Parámetro | Valor |
52
+ |---|---|
53
+ | Clase | `LlamaForCausalLM` (decoder-only, denso, sin MoE) |
54
+ | Parámetros | **88,099,584** (~88M) en `bfloat16` |
55
+ | Capas / hidden / FFN | 12 / 768 / 2048 (SwiGLU) |
56
+ | Atención | GQA 12Q/4KV, head_dim 64, sin bias |
57
+ | Posiciones | RoPE θ=100000, ctx 2048, RMSNorm ε=1e-5 |
58
+ | Embeddings | atados (ahorra ~12.6M) |
59
+ | Archivo | `model.safetensors` (176 MB, 110 tensores, sha `664665e7…`) |
60
+ | Tokenizer | BPE byte-level propio 16k, decoder ByteLevel verificado (~4.0 chars/tok ES/EN) |
61
+ | Especiales | `<pad> <s> </s> <unk> <think> </think>` |
62
 
63
+ ## Entrenamiento
 
 
 
 
 
 
64
 
65
+ 1. **SFT** 58M toks ×2 epochs — loss 0.09, acc 98.4%
66
+ 2. **DPO** β=0.1 — margen de preferencia 4.2
67
+ 3. **GRPO** 550+150 steps, G=4, rewards verificables (formato + fidelidad numérica + anti-overformat)
68
+ 4. Datos: 80M oro humano real degradado (Wikipedia ES/EN, StackExchange, WikiHow) + corpus v4 hasta 625M
69
 
70
+ ## Benchmark (held-out n=1000, greedy)
71
 
72
+ | Global | GFM 0.956 · faith 0.927 · sem 0.888 · tablas 1.0 |
73
+ |---|---|
74
+ | code / ocr / docs / math / tables / html | 0.93 – 1.00 |
75
+ | control (instrucciones estrictas) | 0.19 (limitación conocida) |
76
 
77
+ ## Uso
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78
 
79
+ ```python
80
+ from transformers import AutoModelForCausalLM, AutoTokenizer
81
+ tok = AutoTokenizer.from_pretrained("OpceanAI/PASITA")
82
+ model = AutoModelForCausalLM.from_pretrained("OpceanAI/PASITA", dtype="auto")
83
+ prompt = "CONVIERTE A MARKDOWN:\n" + texto + "\n\n### Markdown:\n"
84
+ ids = tok(prompt, return_tensors="pt", truncation=True, max_length=1024)
85
+ out = model.generate(**ids, max_new_tokens=512, do_sample=False)
86
+ print(tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True))
87
+ ```
88
 
89
+ Régimen válido: documentos medios/largos (OCR, HTML pegado, actas, tutoriales). Frágil en inputs de 1-3 líneas.
90
 
91
+ ## Limitaciones
92
 
93
+ Puede truncar cifras, omitir datos secundarios, emitir H1 ecoicos o continuar tras terminar. Recomendado: `max_new_tokens` adaptativo + beam + rerank por fidelidad.