OpceanAI commited on
Commit
4adddc8
·
verified ·
1 Parent(s): 262104e

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +27 -27
README.md CHANGED
@@ -40,54 +40,54 @@ model-index:
40
  value: 1.0
41
  ---
42
 
43
- # PASITA v1 — plain text → Markdown (fiel)
44
 
45
- Modelo de lenguaje **decoder-only entrenado 100% desde cero** (sin modelo base), especializado en una sola tarea: convertir **texto plano en Markdown válido preservando la información**.
46
 
47
- > Comportamiento de compilador, no de chatbot: añade estructura, no inventa contenido.
48
 
49
- ## Arquitectura (laboratorio)
50
 
51
- | Parámetro | Valor |
52
  |---|---|
53
- | Clase | `LlamaForCausalLM` (decoder-only, denso, sin MoE) |
54
- | Parámetros | **88,099,584** (~88M) en `bfloat16` |
55
- | Capas / hidden / FFN | 12 / 768 / 2048 (SwiGLU) |
56
- | Atención | GQA 12Q/4KV, head_dim 64, sin bias |
57
- | Posiciones | RoPE θ=100000, ctx 2048, RMSNorm ε=1e-5 |
58
- | Embeddings | atados (ahorra ~12.6M) |
59
- | Archivo | `model.safetensors` (176 MB, 110 tensores, sha `664665e7…`) |
60
- | Tokenizer | BPE byte-level propio 16k, decoder ByteLevel verificado (~4.0 chars/tok ES/EN) |
61
- | Especiales | `<pad> <s> </s> <unk> <think> </think>` |
62
 
63
- ## Entrenamiento
64
 
65
- 1. **SFT** 58M toks ×2 epochs — loss 0.09, acc 98.4%
66
- 2. **DPO** β=0.1 — margen de preferencia 4.2
67
- 3. **GRPO** 550+150 steps, G=4, rewards verificables (formato + fidelidad numérica + anti-overformat)
68
- 4. Datos: 80M oro humano real degradado (Wikipedia ES/EN, StackExchange, WikiHow) + corpus v4 hasta 625M
69
 
70
  ## Benchmark (held-out n=1000, greedy)
71
 
72
- | Global | GFM 0.956 · faith 0.927 · sem 0.888 · tablas 1.0 |
73
  |---|---|
74
- | code / ocr / docs / math / tables / html | 0.93 – 1.00 |
75
- | control (instrucciones estrictas) | 0.19 (limitación conocida) |
76
 
77
- ## Uso
78
 
79
  ```python
80
  from transformers import AutoModelForCausalLM, AutoTokenizer
81
  tok = AutoTokenizer.from_pretrained("OpceanAI/PASITA")
82
  model = AutoModelForCausalLM.from_pretrained("OpceanAI/PASITA", dtype="auto")
83
- prompt = "CONVIERTE A MARKDOWN:\n" + texto + "\n\n### Markdown:\n"
84
  ids = tok(prompt, return_tensors="pt", truncation=True, max_length=1024)
85
  out = model.generate(**ids, max_new_tokens=512, do_sample=False)
86
  print(tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True))
87
  ```
88
 
89
- Régimen válido: documentos medios/largos (OCR, HTML pegado, actas, tutoriales). Frágil en inputs de 1-3 líneas.
90
 
91
- ## Limitaciones
92
 
93
- Puede truncar cifras, omitir datos secundarios, emitir H1 ecoicos o continuar tras terminar. Recomendado: `max_new_tokens` adaptativo + beam + rerank por fidelidad.
 
40
  value: 1.0
41
  ---
42
 
43
+ # PASITA v1 — plain text to Markdown (faithful)
44
 
45
+ A **decoder-only language model trained 100% from scratch** (no base model), specialized in a single task: converting **plain text into valid Markdown while preserving information**.
46
 
47
+ > Compiler behavior, not chatbot behavior: it adds structure, it does not invent content.
48
 
49
+ ## Architecture (lab notes)
50
 
51
+ | Parameter | Value |
52
  |---|---|
53
+ | Class | `LlamaForCausalLM` (decoder-only, dense, no MoE) |
54
+ | Parameters | **88,099,584** (~88M) in `bfloat16` |
55
+ | Layers / hidden / FFN | 12 / 768 / 2048 (SwiGLU) |
56
+ | Attention | GQA 12Q/4KV, head_dim 64, no bias |
57
+ | Positions | RoPE theta=100000, ctx 2048, RMSNorm eps=1e-5 |
58
+ | Embeddings | tied (saves ~12.6M params) |
59
+ | File | `model.safetensors` (176 MB, 110 tensors, sha `664665e7…`) |
60
+ | Tokenizer | Custom 16k byte-level BPE, verified ByteLevel decoder (~4.0 chars/token ES/EN) |
61
+ | Special tokens | `<pad> <s> </s> <unk> <think> </think>` |
62
 
63
+ ## Training
64
 
65
+ 1. **SFT** 58M tokens x2 epochs — final loss 0.09, token accuracy 98.4%
66
+ 2. **DPO** beta=0.1 — preference margin 4.2
67
+ 3. **GRPO** 550+150 steps, G=4, verifiable rewards (format + numeric fidelity + anti-overformatting)
68
+ 4. Data: 80M human markdown-derived tokens (ES/EN Wikipedia, StackExchange, WikiHow) scaled to 625M in v4 corpus
69
 
70
  ## Benchmark (held-out n=1000, greedy)
71
 
72
+ | Global | GFM 0.956 - faith 0.927 - sem 0.888 - tables 1.0 |
73
  |---|---|
74
+ | code / ocr / docs / math / tables / html | 0.93 - 1.00 |
75
+ | control (strict instructions) | 0.19 (known limitation) |
76
 
77
+ ## Usage
78
 
79
  ```python
80
  from transformers import AutoModelForCausalLM, AutoTokenizer
81
  tok = AutoTokenizer.from_pretrained("OpceanAI/PASITA")
82
  model = AutoModelForCausalLM.from_pretrained("OpceanAI/PASITA", dtype="auto")
83
+ prompt = "CONVIERTE A MARKDOWN:\n" + text + "\n\n### Markdown:\n"
84
  ids = tok(prompt, return_tensors="pt", truncation=True, max_length=1024)
85
  out = model.generate(**ids, max_new_tokens=512, do_sample=False)
86
  print(tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True))
87
  ```
88
 
89
+ Valid regime: medium/long documents (OCR output, pasted HTML, meeting notes, tutorials). Fragile on 1-3 line inputs.
90
 
91
+ ## Limitations
92
 
93
+ May truncate digits, drop secondary data, emit echo H1s, or continue past completion. Recommended: adaptive `max_new_tokens` + beam search + fidelity rerank.