--- language: - pt license: gemma library_name: gguf pipeline_tag: text-generation base_model: tardellirs/aac-board-generator-770m-ptbr tags: - gguf - llama-cpp - text-generation - gemma3 - aac - augmentative-alternative-communication - arasaac - board-generation - brazilian-portuguese - cpu --- # AAC Board Generator 770M (pt-BR) — GGUF ![Runs on](https://img.shields.io/badge/runs%20on-CPU-success) [![Model](https://img.shields.io/badge/model-card-blue)](https://huggingface.co/tardellirs/aac-board-generator-770m-ptbr) [![License](https://img.shields.io/badge/license-Gemma-lightgrey)](https://ai.google.dev/gemma/terms) [![Used in Papuguinho](https://img.shields.io/badge/used%20in-Papuguinho-ff69b4)](https://www.papuguinho.com) `llama.cpp` **GGUF builds** of [**aac-board-generator-770m-ptbr**](https://huggingface.co/tardellirs/aac-board-generator-770m-ptbr) — a 771.6M Brazilian-Portuguese **multi-function AAC** model that drafts **pictogram-board** word lists *and* follows in-assistant board-editing instructions (add / remove / generate N), on **CPU**. See the [**model card**](https://huggingface.co/tardellirs/aac-board-generator-770m-ptbr) for what it does, training, and evaluation. ## Files | File | Quant | Size | Notes | |---|---|---|---| | `aac-board-generator-770m-ptbr-Q4_K_M.gguf` | Q4_K_M | 561 MB | smallest; slight quality drop | | `aac-board-generator-770m-ptbr-Q5_K_M.gguf` | Q5_K_M | 606 MB | good size/quality balance | | `aac-board-generator-770m-ptbr-Q6_K.gguf` | Q6_K | 766 MB | near-lossless | | `aac-board-generator-770m-ptbr-Q8_0.gguf` | Q8_0 | 824 MB | **recommended** — safe near-lossless | | `aac-board-generator-770m-ptbr-f16.gguf` | f16 | 1.55 GB | full precision reference | At **Q8_0** the model needs on the order of **~0.9–1.1 GB RAM** to serve (weights + a small context) and produces a ~12-item board in **about 4 s on 4 CPU threads**. ## Usage ```bash llama-server -m aac-board-generator-770m-ptbr-Q8_0.gguf -t 4 -c 2048 --host 127.0.0.1 --port 8080 ``` Gemma has no system role — put the instruction in the user turn and call `/completion`: ```python import requests INSTR = ("Você monta pranchas de CAA (pictogramas, pt-BR). Para o PEDIDO, liste ~12 itens concretos e relevantes, " "um por linha, no formato palavra|tipo|sinônimos (tipo: v/s/a/e/l/p). Só a lista.") pedido = "monta uma prancha de brincar no parquinho" prompt = f"user\n{INSTR}\n\nPEDIDO: {pedido}\nmodel\n" print(requests.post("http://127.0.0.1:8080/completion", json={"prompt": prompt, "temperature": 0, "n_predict": 320}).json()["content"]) ``` Generation is **greedy / deterministic** (`temperature 0`). Output is one item per line, `word|type|synonyms`. ## Notes on the build (why these GGUFs are faithful) This model is a **vocabulary-trimmed** Gemma 3 (262k → 64k tokens). A reconstructed SentencePiece tokenizer is *not* bit-exact for a trimmed vocabulary and would silently degrade greedy generation. These GGUFs instead use the **BPE `tokenizer.json`** path, which tokenizes **identically** to the original — so the quantized model reproduces the untrimmed parent's outputs. Verified: identical token counts and identical boards vs. the fp16 reference. ## License & attribution Released under the **[Gemma license](https://ai.google.dev/gemma/terms)** (inherited from Gemma 3). Built for the **ARASAAC** ecosystem (CC BY-NC-SA content). See the [model card](https://huggingface.co/tardellirs/aac-board-generator-770m-ptbr) for full terms. Developed for and used in **Papuguinho** — [www.papuguinho.com](https://www.papuguinho.com).