Text Generation
Transformers
Safetensors
Spanish
English
llama
markdown
text-to-markdown
faithful-generation
from-scratch
spanish
english
ocr-postprocessing
rag
document-understanding
tiny-llm
Eval Results (legacy)
text-generation-inference
Instructions to use OpceanAI/PASITA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OpceanAI/PASITA with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="OpceanAI/PASITA")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("OpceanAI/PASITA") model = AutoModelForCausalLM.from_pretrained("OpceanAI/PASITA", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OpceanAI/PASITA with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OpceanAI/PASITA" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OpceanAI/PASITA", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/OpceanAI/PASITA
- SGLang
How to use OpceanAI/PASITA with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OpceanAI/PASITA" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OpceanAI/PASITA", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OpceanAI/PASITA" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OpceanAI/PASITA", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use OpceanAI/PASITA with Docker Model Runner:
docker model run hf.co/OpceanAI/PASITA
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -3,335 +3,91 @@ license: apache-2.0
|
|
| 3 |
language:
|
| 4 |
- es
|
| 5 |
- en
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
---
|
| 7 |
|
| 8 |
-
PASITA v1
|
| 9 |
|
| 10 |
-
|
| 11 |
|
| 12 |
-
|
| 13 |
|
| 14 |
-
|
| 15 |
|
| 16 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
|
| 18 |
-
|
| 19 |
-
|
|
| 20 |
-
v
|
| 21 |
-
PASITA
|
| 22 |
-
|
|
| 23 |
-
v
|
| 24 |
-
Markdown (GFM)
|
| 25 |
|
| 26 |
-
|
|
|
|
|
|
|
|
|
|
| 27 |
|
| 28 |
-
|
| 29 |
|
| 30 |
-
|
|
|
|
|
|
|
|
|
|
| 31 |
|
| 32 |
-
|
| 33 |
-
Model| PASITA v1
|
| 34 |
-
Task| Plain text → Markdown (GFM)
|
| 35 |
-
Architecture| LlamaForCausalLM
|
| 36 |
-
Parameters| 88,099,584 (~88M)
|
| 37 |
-
Hidden size| 768
|
| 38 |
-
Layers| 12
|
| 39 |
-
Intermediate size| 2048
|
| 40 |
-
Attention heads| 12
|
| 41 |
-
Key/value heads| 4
|
| 42 |
-
Attention| Grouped Query Attention (GQA)
|
| 43 |
-
Head dimension| 64
|
| 44 |
-
Activation| SiLU / SwiGLU
|
| 45 |
-
Normalization| RMSNorm
|
| 46 |
-
RMSNorm epsilon| 1e-5
|
| 47 |
-
Positional encoding| RoPE
|
| 48 |
-
RoPE theta| 100,000
|
| 49 |
-
Maximum context| 2,048 tokens
|
| 50 |
-
Vocabulary| 16,384 tokens
|
| 51 |
-
Word embeddings| Tied
|
| 52 |
-
Attention bias| Disabled
|
| 53 |
-
MLP bias| Disabled
|
| 54 |
-
Parameter dtype| bfloat16
|
| 55 |
-
Weight format| safetensors
|
| 56 |
-
Number of tensors| 110
|
| 57 |
|
| 58 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 59 |
|
| 60 |
-
|
| 61 |
|
| 62 |
-
|
| 63 |
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
- RMSNorm
|
| 67 |
-
- SiLU activation
|
| 68 |
-
- SwiGLU feed-forward blocks
|
| 69 |
-
- Rotary Position Embeddings (RoPE)
|
| 70 |
-
- A maximum sequence length of 2,048 tokens
|
| 71 |
-
- Tied input and output embeddings
|
| 72 |
-
- No attention or MLP biases
|
| 73 |
-
|
| 74 |
-
The model is implemented using "LlamaForCausalLM".
|
| 75 |
-
|
| 76 |
-
Tokenizer
|
| 77 |
-
|
| 78 |
-
PASITA uses a custom Byte-Level BPE tokenizer with a vocabulary of 16,384 pieces.
|
| 79 |
-
|
| 80 |
-
The tokenizer includes:
|
| 81 |
-
|
| 82 |
-
<pad>
|
| 83 |
-
<s>
|
| 84 |
-
</s>
|
| 85 |
-
<unk>
|
| 86 |
-
<think>
|
| 87 |
-
</think>
|
| 88 |
-
|
| 89 |
-
The ByteLevel pre-tokenizer and decoder have been verified using exact round-trip tests.
|
| 90 |
-
|
| 91 |
-
Measured tokenization efficiency is approximately:
|
| 92 |
-
|
| 93 |
-
- 4.0 characters per token for Spanish and English
|
| 94 |
-
- 2.0 characters per token for code
|
| 95 |
-
|
| 96 |
-
Training Data
|
| 97 |
-
|
| 98 |
-
PASITA was trained on approximately 80 million BPE tokens.
|
| 99 |
-
|
| 100 |
-
The underlying corpus consists of human-produced Markdown obtained from public sources, including:
|
| 101 |
-
|
| 102 |
-
- GoodWiki English
|
| 103 |
-
- Spanish Wikipedia dumps
|
| 104 |
-
- StackExchange Q&A
|
| 105 |
-
- WikiHow
|
| 106 |
-
- Cosmopedia
|
| 107 |
-
|
| 108 |
-
The Markdown documents were deterministically degraded to produce noisy plain-text inputs.
|
| 109 |
-
|
| 110 |
-
The resulting training data contains approximately 37% Spanish and 63% English.
|
| 111 |
-
|
| 112 |
-
Dataset Splits
|
| 113 |
-
|
| 114 |
-
Split| Tokens / prompts| Description
|
| 115 |
-
SFT| 58M / 47k| Supervised fine-tuning
|
| 116 |
-
DPO| 16M / 9k pairs| Preference optimization
|
| 117 |
-
GRPO| 6M / 10.4k prompts| Group-relative optimization
|
| 118 |
-
Benchmark| 2.5M / 2k| Held-out evaluation
|
| 119 |
-
|
| 120 |
-
The benchmark split is disjoint from the training data by hash.
|
| 121 |
-
|
| 122 |
-
Dataset quality controls included Markdown validity checks, hash and ROUGE-based deduplication, context-length validation, and multi-judge verification.
|
| 123 |
-
|
| 124 |
-
Training
|
| 125 |
-
|
| 126 |
-
PASITA v1 was trained using bfloat16 on NVIDIA B300 and A100 hardware.
|
| 127 |
-
|
| 128 |
-
The training pipeline consisted of three stages.
|
| 129 |
-
|
| 130 |
-
1. Supervised Fine-Tuning
|
| 131 |
-
|
| 132 |
-
SFT was performed for two epochs.
|
| 133 |
-
|
| 134 |
-
Reported final metrics:
|
| 135 |
-
|
| 136 |
-
- Loss: approximately 0.09
|
| 137 |
-
- Token accuracy: 98.4%
|
| 138 |
-
|
| 139 |
-
2. Direct Preference Optimization
|
| 140 |
-
|
| 141 |
-
DPO was performed for one epoch with:
|
| 142 |
-
|
| 143 |
-
β = 0.1
|
| 144 |
-
|
| 145 |
-
The reported final preference margin was approximately 4.2.
|
| 146 |
-
|
| 147 |
-
The original DPO run contained partially spurious signals in some rejected examples. A subsequent DPO-v2 dataset was developed with more minimal edits and harder negatives.
|
| 148 |
-
|
| 149 |
-
3. Group Relative Policy Optimization
|
| 150 |
-
|
| 151 |
-
GRPO was performed for 550 steps using groups of four generations.
|
| 152 |
-
|
| 153 |
-
The reward design focused primarily on structural correctness and content fidelity.
|
| 154 |
-
|
| 155 |
-
During the original training process, the fidelity component of the reward exhibited noise related to the decoder configuration, while structural rewards were learned more consistently.
|
| 156 |
-
|
| 157 |
-
Evaluation
|
| 158 |
-
|
| 159 |
-
Evaluation was performed greedily on held-out data.
|
| 160 |
-
|
| 161 |
-
Bench-500
|
| 162 |
-
|
| 163 |
-
Metric| Score
|
| 164 |
-
GFM validity| 0.958
|
| 165 |
-
Faithfulness| 0.930
|
| 166 |
-
Semantic faithfulness| 0.880
|
| 167 |
-
Table correctness| 1.000
|
| 168 |
-
|
| 169 |
-
Bench-1000
|
| 170 |
-
|
| 171 |
-
Metric| Score
|
| 172 |
-
GFM validity| 0.956
|
| 173 |
-
Faithfulness| 0.927
|
| 174 |
-
Semantic faithfulness| 0.888
|
| 175 |
-
Table correctness| 1.000
|
| 176 |
-
|
| 177 |
-
Domain Results
|
| 178 |
-
|
| 179 |
-
Domain| Score
|
| 180 |
-
Code| 0.965
|
| 181 |
-
OCR / transcripts| 0.953
|
| 182 |
-
Documents| 0.946
|
| 183 |
-
Mathematics| 0.981
|
| 184 |
-
Tables| 0.945
|
| 185 |
-
HTML| 1.000
|
| 186 |
-
Control| 0.190
|
| 187 |
-
|
| 188 |
-
The control category is the principal known weakness of the current version.
|
| 189 |
-
|
| 190 |
-
Inference
|
| 191 |
-
|
| 192 |
-
PASITA expects a prompt beginning with "CONVIERTE A MARKDOWN", followed by the source text and the "### Markdown:" separator.
|
| 193 |
-
|
| 194 |
-
Example:
|
| 195 |
-
|
| 196 |
-
prompt = (
|
| 197 |
-
"CONVIERTE A MARKDOWN:\n"
|
| 198 |
-
+ text
|
| 199 |
-
+ "\n\n### Markdown:\n"
|
| 200 |
-
)
|
| 201 |
-
|
| 202 |
-
A recommended deterministic generation configuration is:
|
| 203 |
-
|
| 204 |
-
generate(
|
| 205 |
-
do_sample=False,
|
| 206 |
-
max_new_tokens=int(input_length * 1.3 + 128)
|
| 207 |
-
)
|
| 208 |
-
|
| 209 |
-
The model is intended for deterministic transformation rather than conversational generation.
|
| 210 |
-
|
| 211 |
-
Model Size and Runtime
|
| 212 |
-
|
| 213 |
-
The model contains approximately 88 million parameters.
|
| 214 |
-
|
| 215 |
-
The stored model weights occupy approximately:
|
| 216 |
-
|
| 217 |
-
176,211,400 bytes
|
| 218 |
-
|
| 219 |
-
or approximately 168 MB.
|
| 220 |
-
|
| 221 |
-
The weights are stored natively in the safetensors format.
|
| 222 |
-
|
| 223 |
-
The model uses bfloat16 parameters, while CPU inference can be performed using a float32-loaded model when required by the runtime configuration.
|
| 224 |
-
|
| 225 |
-
Approximate memory usage for the weights is around 180 MB, excluding runtime allocations and KV-cache requirements.
|
| 226 |
-
|
| 227 |
-
The model is therefore small enough to be deployed on CPU systems and on GPUs with at least 2 GB of memory, subject to runtime overhead and the selected inference configuration.
|
| 228 |
-
|
| 229 |
-
Weight Integrity
|
| 230 |
-
|
| 231 |
-
The primary model file is:
|
| 232 |
-
|
| 233 |
-
final_model/model.safetensors
|
| 234 |
-
|
| 235 |
-
SHA-256 prefix:
|
| 236 |
-
|
| 237 |
-
664665e7ee2b0002
|
| 238 |
-
|
| 239 |
-
Full hashes should be obtained from the corresponding release artifacts when verifying a downloaded model.
|
| 240 |
-
|
| 241 |
-
Intended Use
|
| 242 |
-
|
| 243 |
-
PASITA is designed for applications that require deterministic or near-deterministic conversion of unstructured text into Markdown.
|
| 244 |
-
|
| 245 |
-
Potential applications include:
|
| 246 |
-
|
| 247 |
-
- Plain-text document formatting
|
| 248 |
-
- Markdown normalization pipelines
|
| 249 |
-
- Documentation preprocessing
|
| 250 |
-
- Text-to-Markdown conversion
|
| 251 |
-
- Local document transformation
|
| 252 |
-
- CPU-based Markdown processing
|
| 253 |
-
|
| 254 |
-
PASITA is not designed or evaluated as a general-purpose conversational assistant.
|
| 255 |
-
|
| 256 |
-
Known Limitations
|
| 257 |
-
|
| 258 |
-
PASITA v1 has several known limitations.
|
| 259 |
-
|
| 260 |
-
Content Preservation
|
| 261 |
-
|
| 262 |
-
Although faithfulness is substantially stronger than its Markdown validity score alone would indicate, the model can still:
|
| 263 |
-
|
| 264 |
-
- Truncate numerical values
|
| 265 |
-
- Omit secondary information
|
| 266 |
-
- Echo headings unnecessarily
|
| 267 |
-
- Continue generating after completing the transformation
|
| 268 |
-
|
| 269 |
-
For example, numerical content such as "23.3" may occasionally be reduced to "23".
|
| 270 |
-
|
| 271 |
-
Short Inputs
|
| 272 |
-
|
| 273 |
-
The current model is less reliable on very short inputs, particularly inputs consisting of only one to three lines.
|
| 274 |
-
|
| 275 |
-
Strict Formatting Instructions
|
| 276 |
-
|
| 277 |
-
The control benchmark remains substantially weaker than the other evaluated domains. PASITA should therefore not be assumed to reliably satisfy arbitrary, highly specific formatting constraints.
|
| 278 |
-
|
| 279 |
-
Generation Termination
|
| 280 |
-
|
| 281 |
-
Under greedy generation, PASITA can occasionally continue generating after the correct Markdown transformation has already been produced.
|
| 282 |
-
|
| 283 |
-
Applications should therefore use an appropriate generation limit and explicitly handle end-of-sequence behavior.
|
| 284 |
-
|
| 285 |
-
License
|
| 286 |
-
|
| 287 |
-
PASITA v1 is released under the Apache License 2.0.
|
| 288 |
-
|
| 289 |
-
Copyright © 2026 OpceanAI.
|
| 290 |
-
|
| 291 |
-
The Apache License 2.0 permits use, reproduction, modification, distribution, and creation of derivative works subject to the terms and conditions of the license.
|
| 292 |
-
|
| 293 |
-
The complete license text is provided in the ""LICENSE"" (LICENSE) file.
|
| 294 |
-
|
| 295 |
-
The Apache License 2.0 applies to the PASITA project and its released model artifacts. Third-party datasets and source material used during training may be subject to separate licenses, attribution requirements, or other terms.
|
| 296 |
-
|
| 297 |
-
Users intending to redistribute or commercially deploy systems trained on the underlying datasets should independently review the applicable licenses and attribution requirements, including the requirements associated with Wikipedia-derived material.
|
| 298 |
-
|
| 299 |
-
Data Provenance
|
| 300 |
-
|
| 301 |
-
The model weights were produced from a random initialization and were not initialized from a pretrained language model.
|
| 302 |
-
|
| 303 |
-
Training data was derived from publicly available sources, including GoodWiki, Wikipedia, StackExchange, WikiHow, and Cosmopedia.
|
| 304 |
-
|
| 305 |
-
The licensing status of the resulting model does not supersede the licensing requirements applicable to third-party source material.
|
| 306 |
-
|
| 307 |
-
Project Status
|
| 308 |
-
|
| 309 |
-
PASITA v1 should be considered an experimental specialized language model.
|
| 310 |
-
|
| 311 |
-
The current results demonstrate that a compact model trained specifically for plain-text-to-Markdown transformation can achieve high content-fidelity scores on held-out evaluation data while remaining small enough for practical local inference.
|
| 312 |
-
|
| 313 |
-
Future development is focused on improving:
|
| 314 |
-
|
| 315 |
-
- Termination behavior
|
| 316 |
-
- Control and instruction-following performance
|
| 317 |
-
- Content preservation
|
| 318 |
-
- Robustness on short inputs
|
| 319 |
-
- Markdown structural correctness
|
| 320 |
-
- Preference and reinforcement-learning data quality
|
| 321 |
-
|
| 322 |
-
Summary
|
| 323 |
-
|
| 324 |
-
PASITA v1 is an approximately 88M-parameter decoder-only Transformer trained from scratch specifically for plain-text-to-Markdown transformation.
|
| 325 |
-
|
| 326 |
-
Its current Bench-1000 results are:
|
| 327 |
-
|
| 328 |
-
GFM validity 95.6%
|
| 329 |
-
Faithfulness 92.7%
|
| 330 |
-
Semantic faith. 88.8%
|
| 331 |
-
Table correctness 100.0%
|
| 332 |
-
|
| 333 |
-
The model's primary strength is its ability to perform structured Markdown transformation while preserving most of the source information.
|
| 334 |
-
|
| 335 |
-
Its primary weaknesses are strict control tasks, very short inputs, and occasional post-completion generation.
|
| 336 |
-
|
| 337 |
-
PASITA is therefore best understood as a small, specialized text-transformation model, rather than a general-purpose LLM.
|
|
|
|
| 3 |
language:
|
| 4 |
- es
|
| 5 |
- en
|
| 6 |
+
library_name: transformers
|
| 7 |
+
pipeline_tag: text-generation
|
| 8 |
+
tags:
|
| 9 |
+
- markdown
|
| 10 |
+
- text-to-markdown
|
| 11 |
+
- faithful-generation
|
| 12 |
+
- llama
|
| 13 |
+
- from-scratch
|
| 14 |
+
- spanish
|
| 15 |
+
- english
|
| 16 |
+
- ocr-postprocessing
|
| 17 |
+
- rag
|
| 18 |
+
- document-understanding
|
| 19 |
+
- tiny-llm
|
| 20 |
+
model-index:
|
| 21 |
+
- name: PASITA
|
| 22 |
+
results:
|
| 23 |
+
- task:
|
| 24 |
+
type: text-generation
|
| 25 |
+
dataset:
|
| 26 |
+
name: PASITA-bench-1000 (held-out, private)
|
| 27 |
+
type: custom
|
| 28 |
+
metrics:
|
| 29 |
+
- name: gfm_validity
|
| 30 |
+
type: accuracy
|
| 31 |
+
value: 0.956
|
| 32 |
+
- name: faithfulness
|
| 33 |
+
type: faithfulness
|
| 34 |
+
value: 0.927
|
| 35 |
+
- name: semantic_faithfulness
|
| 36 |
+
type: faithfulness
|
| 37 |
+
value: 0.888
|
| 38 |
+
- name: table_fidelity
|
| 39 |
+
type: accuracy
|
| 40 |
+
value: 1.0
|
| 41 |
---
|
| 42 |
|
| 43 |
+
# PASITA v1 — plain text → Markdown (fiel)
|
| 44 |
|
| 45 |
+
Modelo de lenguaje **decoder-only entrenado 100% desde cero** (sin modelo base), especializado en una sola tarea: convertir **texto plano en Markdown válido preservando la información**.
|
| 46 |
|
| 47 |
+
> Comportamiento de compilador, no de chatbot: añade estructura, no inventa contenido.
|
| 48 |
|
| 49 |
+
## Arquitectura (laboratorio)
|
| 50 |
|
| 51 |
+
| Parámetro | Valor |
|
| 52 |
+
|---|---|
|
| 53 |
+
| Clase | `LlamaForCausalLM` (decoder-only, denso, sin MoE) |
|
| 54 |
+
| Parámetros | **88,099,584** (~88M) en `bfloat16` |
|
| 55 |
+
| Capas / hidden / FFN | 12 / 768 / 2048 (SwiGLU) |
|
| 56 |
+
| Atención | GQA 12Q/4KV, head_dim 64, sin bias |
|
| 57 |
+
| Posiciones | RoPE θ=100000, ctx 2048, RMSNorm ε=1e-5 |
|
| 58 |
+
| Embeddings | atados (ahorra ~12.6M) |
|
| 59 |
+
| Archivo | `model.safetensors` (176 MB, 110 tensores, sha `664665e7…`) |
|
| 60 |
+
| Tokenizer | BPE byte-level propio 16k, decoder ByteLevel verificado (~4.0 chars/tok ES/EN) |
|
| 61 |
+
| Especiales | `<pad> <s> </s> <unk> <think> </think>` |
|
| 62 |
|
| 63 |
+
## Entrenamiento
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 64 |
|
| 65 |
+
1. **SFT** 58M toks ×2 epochs — loss 0.09, acc 98.4%
|
| 66 |
+
2. **DPO** β=0.1 — margen de preferencia 4.2
|
| 67 |
+
3. **GRPO** 550+150 steps, G=4, rewards verificables (formato + fidelidad numérica + anti-overformat)
|
| 68 |
+
4. Datos: 80M oro humano real degradado (Wikipedia ES/EN, StackExchange, WikiHow) + corpus v4 hasta 625M
|
| 69 |
|
| 70 |
+
## Benchmark (held-out n=1000, greedy)
|
| 71 |
|
| 72 |
+
| Global | GFM 0.956 · faith 0.927 · sem 0.888 · tablas 1.0 |
|
| 73 |
+
|---|---|
|
| 74 |
+
| code / ocr / docs / math / tables / html | 0.93 – 1.00 |
|
| 75 |
+
| control (instrucciones estrictas) | 0.19 (limitación conocida) |
|
| 76 |
|
| 77 |
+
## Uso
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
+
```python
|
| 80 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 81 |
+
tok = AutoTokenizer.from_pretrained("OpceanAI/PASITA")
|
| 82 |
+
model = AutoModelForCausalLM.from_pretrained("OpceanAI/PASITA", dtype="auto")
|
| 83 |
+
prompt = "CONVIERTE A MARKDOWN:\n" + texto + "\n\n### Markdown:\n"
|
| 84 |
+
ids = tok(prompt, return_tensors="pt", truncation=True, max_length=1024)
|
| 85 |
+
out = model.generate(**ids, max_new_tokens=512, do_sample=False)
|
| 86 |
+
print(tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True))
|
| 87 |
+
```
|
| 88 |
|
| 89 |
+
Régimen válido: documentos medios/largos (OCR, HTML pegado, actas, tutoriales). Frágil en inputs de 1-3 líneas.
|
| 90 |
|
| 91 |
+
## Limitaciones
|
| 92 |
|
| 93 |
+
Puede truncar cifras, omitir datos secundarios, emitir H1 ecoicos o continuar tras terminar. Recomendado: `max_new_tokens` adaptativo + beam + rerank por fidelidad.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|