--- license: apache-2.0 base_model: Qwen/Qwen2.5-0.5B-Instruct tags: - gguf - qwen2.5 - 0.5b - fine-tuned - structured-output - json - intent-classification - slot-filling - on-device - personal-finance - multilingual - llama-cpp language: - en - it - es - fr pipeline_tag: text-generation --- LoRA fine-tune of Qwen2.5 **0.5B** (~398 MB Q4_K_M) for phones: small enough to download over cellular and run on-device, specialized for personal-finance intent JSON (EN/IT/ES/FR). # AiBudget Intent 0.5B v5 **On-device personal-finance intent → strict JSON** (LoRA SFT on Qwen2.5-0.5B-Instruct, fused, **Q4_K_M GGUF**) > Maps natural-language questions about personal spending (EN / IT / ES / FR) into a single, schema-constrained JSON intent object. > The AiBudget mobile app parses the JSON and runs **local SQLite** tools. When users choose this on-device model, **transaction data is not sent to a cloud LLM for intent parsing**. **Why a 0.5B model?** We deliberately started from a **very small** instruct model so the shipped artifact stays around **one download (~400 MB)** and can run **locally on mid-range phones** without a cloud API. LoRA SFT narrows that base model to one job—structured budget intents—not open-ended chat. Larger models can be smarter in general, but they are harder to justify for privacy-first, offline-friendly mobile installs. **Published artifact:** fused, quantized **GGUF only** — LoRA adapter weights are **not** in this repository. **Recommended download:** [`aibudget-intent-05b-v5-Q4_K_M.gguf`](https://huggingface.co/pablitoelpeligro/aibudget-intent-05b-GGUF/resolve/main/aibudget-intent-05b-v5-Q4_K_M.gguf) (~398 MB, SHA256 `d3d6fc0780359cd4d198e9354b9ecfc0ff9cfb0fb729f040326fa6e2f02c693a`) Other `aibudget-intent-05b-v*-Q4_K_M.gguf` files in this repo are **experimental rounds**; use **v5** unless you are reproducing eval history. --- ## Model details | | | |---|---| | **Developed by** | pablitoelpeligro (AiBudget) | | **Model type** | Causal LM; LoRA SFT → merged weights → GGUF Q4_K_M | | **Languages** | English, Italian, Spanish, French | | **License** | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) (same as base) | | **Base model** | [Qwen/Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct) | | **Training** | [mlx-lm](https://github.com/ml-explore/mlx-examples/tree/main/llms/mlx_lm) on Apple Silicon → fuse → [llama.cpp](https://github.com/ggml-org/llama.cpp) `Q4_K_M` | | **Production inference** | [llama.rn](https://github.com/mybigday/llama.rn) (React Native) + **GBNF** grammar matching the intent JSON schema | --- ## Intended use ### Direct use Drop-in **intent parser** for personal-finance assistants (especially mobile / edge): - Total spend, category spend, merchant (`description`) spend - Transaction search, sorting, limits - Budget status, monthly trends, top merchants - Period comparisons (`compare_periods`) - Structured `add_transaction` extraction - `unknown` for meta / out-of-scope questions Trained and evaluated with the same **`compact` prompt profile** as production (short system text, **no** few-shot examples in the prompt). ### Out-of-scope - General open-ended chat - Non–personal-finance domains - Financial advice without human review - Languages outside EN / IT / ES / FR --- ## How to use ### Download ```bash hf download pablitoelpeligro/aibudget-intent-05b-GGUF \ aibudget-intent-05b-v5-Q4_K_M.gguf --local-dir . ``` ### Prompt format (compact) Production builds chat messages like this (Qwen chat template): 1. **System:** compact classifier text + newline + `conversationContext: {"today":"…",…}` (JSON) 2. **Optional:** recent **user** turns only (no assistant prose in history) 3. **User:** current question Example system line (English app, English query): ```text Translate the user's message into exactly one AiBudget financial intent. Use conversationContext and earlier user messages to resolve follow-ups. Message language: en. App language: en. Return only JSON. conversationContext: {"today":"2026-09-22","currency":"USD"} ``` Use **`temperature: 0`** for reproducible decoding (matches our offline eval). ### llama.cpp (best-effort without grammar) Without **GBNF**, the model may still emit JSON but slot accuracy can drop vs the app. For a quick smoke test: ```bash # Flag names vary by llama.cpp version; use Qwen2.5-Instruct chat template in production. llama-cli -m aibudget-intent-05b-v5-Q4_K_M.gguf --temp 0 -n 512 \ -sys "Translate the user's message into exactly one AiBudget financial intent. Return only JSON." \ -p "Total spending this month" ``` ### AiBudget app In **Settings**, download **AiBudget Intent 0.5B (tuned v5)**. The app loads this GGUF with llama.rn and **grammar-constrained** decoding so output matches the schema below. --- ## Output schema One JSON object per turn. **All keys required**; unused slots are `null`. `intent` and `dateRange` are always strings (never null). `confidence` is in `[0, 1]`. | Field | Type | Role | |-------|------|------| | `intent` | enum | `total_spend` · `description_spend` · `category_spend` · `category_breakdown` · `monthly_trends` · `budget_status` · `transaction_search` · `add_transaction` · `net_cashflow` · `compare_periods` · `top_descriptions` · `unknown` | | `description` | string \| null | Merchant / payee (1–4 words) | | `category` | string \| null | Category label | | `dateRange` | enum | `this_month` · `last_month` · `today` · `yesterday` · `previous_range` · `last_n_months` · `this_year` · `last_year` · `year_to_date` · `quarter` · `custom` · `all_time` · `unspecified` | | `monthsCount` | int \| null | For `last_n_months` | | `startDate`, `endDate` | string \| null | Custom range when `dateRange` is `custom` | | `periodLabel` | string \| null | Free-text period hint | | `amount` | number \| null | e.g. search by amount | | `notes` | string \| null | e.g. `add_transaction` | | `sort` | enum \| null | `amount_desc` · `amount_asc` · `date_desc` · `date_asc` | | `limit` | int \| null | Row cap for lists / tops | | `metric` | enum \| null | `sum` · `max` · `min` · `count` | | `transactionType` | enum \| null | `expense` · `income` | | `ordinal` | int \| null | 1 = most recent in sort order | | `offset` | int \| null | Pagination offset | | `quarterNumber` | int \| null | 1–4 when `dateRange` is `quarter` | | `year` | int \| null | Year for quarter / year scopes | | `dateRangeA`, `dateRangeB` | enum \| null | Same values as `dateRange`; period comparison | | `startDateB`, `endDateB`, `periodLabelB` | string \| null | Second period for custom compare | | `groupBy` | enum \| null | `category` · `description` | | `compareMetric` | enum \| null | `sum` · `delta` · `pct_change` | | `confidence` | number | Model confidence | ### Example **User:** `how much did i blow on Uber last month` ```json { "intent": "description_spend", "description": "Uber", "category": null, "dateRange": "last_month", "monthsCount": null, "startDate": null, "endDate": null, "periodLabel": null, "amount": null, "notes": null, "sort": null, "limit": null, "metric": null, "transactionType": null, "ordinal": null, "offset": null, "quarterNumber": null, "year": null, "dateRangeA": null, "dateRangeB": null, "startDateB": null, "endDateB": null, "periodLabelB": null, "groupBy": null, "compareMetric": null, "confidence": 0.9 } ``` --- ## Training | Setting | Value | |---------|--------| | Framework | mlx-lm (`mlx_lm.lora`) | | Base | `Qwen/Qwen2.5-0.5B-Instruct` (full precision, not QLoRA) | | Method | LoRA rank **8**, 16 layers, dropout 0.05, MLX scale **20**, lr **2e-5** | | Steps | **1200** iterations, batch 4 × grad accum 2, max seq **1536** | | Loss | **Response only** (`mask_prompt: true`) | | Export | Fuse adapter → GGUF F16 → **Q4_K_M** (evaluated artifact) | | Hardware | Apple Silicon (MLX) | **Data (not on Hugging Face):** ~800 hand-written seed rows (EN/IT/ES/FR), ~83 curated examples, ~2k synthetic LLM-generated rows (schema- and leak-checked), plus small targeted rounds (v5 added ~115 rows for unknown / category narrows). Typical MLX split: ~**2.5k–2.8k** train / ~**280–310** valid. Training chat format matches production (`compact` system prompt, no in-prompt few-shots). --- ## Evaluation Offline eval on a **75-prompt holdout** (held out from training) and a **310-prompt dev** stress set. Metrics are measured on the **quantized GGUF** at **temperature 0**, **without** post-hoc regex fixes on model output (raw model JSON only). ### Holdout (primary) | Metric | Stock Qwen2.5-0.5B Q4_K_M | **This model (v5)** | |--------|---------------------------|---------------------| | Task success | 50.7% | **94.7%** | | Confidently wrong | 34.7% | **4.0%** | | Intent match | — | 82.7% | | Full JSON match (intent + slots) | — | 64.0% | | p50 / p90 latency (Mac eval) | — | ~499 ms / ~515 ms | Release bar on holdout: task success ≥ 80%, confidently wrong ≤ 8% — **met** for v5. ### Dev (internal, harder) Task success **~70%**, confidently wrong **~26%** — includes multi-turn and edge cases; not the primary release gate. **Caveats:** Eval scores JSON + routing, not final UX copy. Some merchant→category follow-ups remain imperfect. Quantization may differ slightly from fused F16. --- ## Limitations - **Domain-specific:** personal finance phrasing only. - **Schema-bound:** best results with GBNF or strict JSON validation. - **No PEFT files:** you cannot merge a published LoRA from this repo; only the fused GGUF is available. - **Training data:** synthetic + curated; not a public dataset on the Hub. - **Not** a replacement for licensed financial advice. --- ## License Apache-2.0, consistent with [Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct). Use under the same terms as the base model and respect Qwen usage policies. --- ## Citation / links - **This repository:** https://huggingface.co/pablitoelpeligro/aibudget-intent-05b-GGUF - **Base model:** https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct If you use this model, please cite the base Qwen model and link to this GGUF repository.