LoRA fine-tune of Qwen2.5 0.5B (~398 MB Q4_K_M) for phones: small enough to download over cellular and run on-device, specialized for personal-finance intent JSON (EN/IT/ES/FR).

AiBudget Intent 0.5B v5

On-device personal-finance intent โ†’ strict JSON (LoRA SFT on Qwen2.5-0.5B-Instruct, fused, Q4_K_M GGUF)

Maps natural-language questions about personal spending (EN / IT / ES / FR) into a single, schema-constrained JSON intent object.
The AiBudget mobile app parses the JSON and runs local SQLite tools. When users choose this on-device model, transaction data is not sent to a cloud LLM for intent parsing.

Why a 0.5B model? We deliberately started from a very small instruct model so the shipped artifact stays around one download (~400 MB) and can run locally on mid-range phones without a cloud API. LoRA SFT narrows that base model to one jobโ€”structured budget intentsโ€”not open-ended chat. Larger models can be smarter in general, but they are harder to justify for privacy-first, offline-friendly mobile installs.

Published artifact: fused, quantized GGUF only โ€” LoRA adapter weights are not in this repository.

Recommended download: aibudget-intent-05b-v5-Q4_K_M.gguf (~398 MB, SHA256 d3d6fc0780359cd4d198e9354b9ecfc0ff9cfb0fb729f040326fa6e2f02c693a)

Other aibudget-intent-05b-v*-Q4_K_M.gguf files in this repo are experimental rounds; use v5 unless you are reproducing eval history.


Model details

Developed by pablitoelpeligro (AiBudget)
Model type Causal LM; LoRA SFT โ†’ merged weights โ†’ GGUF Q4_K_M
Languages English, Italian, Spanish, French
License Apache-2.0 (same as base)
Base model Qwen/Qwen2.5-0.5B-Instruct
Training mlx-lm on Apple Silicon โ†’ fuse โ†’ llama.cpp Q4_K_M
Production inference llama.rn (React Native) + GBNF grammar matching the intent JSON schema

Intended use

Direct use

Drop-in intent parser for personal-finance assistants (especially mobile / edge):

  • Total spend, category spend, merchant (description) spend
  • Transaction search, sorting, limits
  • Budget status, monthly trends, top merchants
  • Period comparisons (compare_periods)
  • Structured add_transaction extraction
  • unknown for meta / out-of-scope questions

Trained and evaluated with the same compact prompt profile as production (short system text, no few-shot examples in the prompt).

Out-of-scope

  • General open-ended chat
  • Nonโ€“personal-finance domains
  • Financial advice without human review
  • Languages outside EN / IT / ES / FR

How to use

Download

hf download pablitoelpeligro/aibudget-intent-05b-GGUF \
  aibudget-intent-05b-v5-Q4_K_M.gguf --local-dir .

Prompt format (compact)

Production builds chat messages like this (Qwen chat template):

  1. System: compact classifier text + newline + conversationContext: {"today":"โ€ฆ",โ€ฆ} (JSON)
  2. Optional: recent user turns only (no assistant prose in history)
  3. User: current question

Example system line (English app, English query):

Translate the user's message into exactly one AiBudget financial intent. Use conversationContext and earlier user messages to resolve follow-ups. Message language: en. App language: en. Return only JSON.
conversationContext: {"today":"2026-09-22","currency":"USD"}

Use temperature: 0 for reproducible decoding (matches our offline eval).

llama.cpp (best-effort without grammar)

Without GBNF, the model may still emit JSON but slot accuracy can drop vs the app. For a quick smoke test:

# Flag names vary by llama.cpp version; use Qwen2.5-Instruct chat template in production.
llama-cli -m aibudget-intent-05b-v5-Q4_K_M.gguf --temp 0 -n 512 \
  -sys "Translate the user's message into exactly one AiBudget financial intent. Return only JSON." \
  -p "Total spending this month"

AiBudget app

In Settings, download AiBudget Intent 0.5B (tuned v5). The app loads this GGUF with llama.rn and grammar-constrained decoding so output matches the schema below.


Output schema

One JSON object per turn. All keys required; unused slots are null. intent and dateRange are always strings (never null). confidence is in [0, 1].

Field Type Role
intent enum total_spend ยท description_spend ยท category_spend ยท category_breakdown ยท monthly_trends ยท budget_status ยท transaction_search ยท add_transaction ยท net_cashflow ยท compare_periods ยท top_descriptions ยท unknown
description string | null Merchant / payee (1โ€“4 words)
category string | null Category label
dateRange enum this_month ยท last_month ยท today ยท yesterday ยท previous_range ยท last_n_months ยท this_year ยท last_year ยท year_to_date ยท quarter ยท custom ยท all_time ยท unspecified
monthsCount int | null For last_n_months
startDate, endDate string | null Custom range when dateRange is custom
periodLabel string | null Free-text period hint
amount number | null e.g. search by amount
notes string | null e.g. add_transaction
sort enum | null amount_desc ยท amount_asc ยท date_desc ยท date_asc
limit int | null Row cap for lists / tops
metric enum | null sum ยท max ยท min ยท count
transactionType enum | null expense ยท income
ordinal int | null 1 = most recent in sort order
offset int | null Pagination offset
quarterNumber int | null 1โ€“4 when dateRange is quarter
year int | null Year for quarter / year scopes
dateRangeA, dateRangeB enum | null Same values as dateRange; period comparison
startDateB, endDateB, periodLabelB string | null Second period for custom compare
groupBy enum | null category ยท description
compareMetric enum | null sum ยท delta ยท pct_change
confidence number Model confidence

Example

User: how much did i blow on Uber last month

{
  "intent": "description_spend",
  "description": "Uber",
  "category": null,
  "dateRange": "last_month",
  "monthsCount": null,
  "startDate": null,
  "endDate": null,
  "periodLabel": null,
  "amount": null,
  "notes": null,
  "sort": null,
  "limit": null,
  "metric": null,
  "transactionType": null,
  "ordinal": null,
  "offset": null,
  "quarterNumber": null,
  "year": null,
  "dateRangeA": null,
  "dateRangeB": null,
  "startDateB": null,
  "endDateB": null,
  "periodLabelB": null,
  "groupBy": null,
  "compareMetric": null,
  "confidence": 0.9
}

Training

Setting Value
Framework mlx-lm (mlx_lm.lora)
Base Qwen/Qwen2.5-0.5B-Instruct (full precision, not QLoRA)
Method LoRA rank 8, 16 layers, dropout 0.05, MLX scale 20, lr 2e-5
Steps 1200 iterations, batch 4 ร— grad accum 2, max seq 1536
Loss Response only (mask_prompt: true)
Export Fuse adapter โ†’ GGUF F16 โ†’ Q4_K_M (evaluated artifact)
Hardware Apple Silicon (MLX)

Data (not on Hugging Face): ~800 hand-written seed rows (EN/IT/ES/FR), ~83 curated examples, ~2k synthetic LLM-generated rows (schema- and leak-checked), plus small targeted rounds (v5 added ~115 rows for unknown / category narrows). Typical MLX split: ~2.5kโ€“2.8k train / ~280โ€“310 valid. Training chat format matches production (compact system prompt, no in-prompt few-shots).


Evaluation

Offline eval on a 75-prompt holdout (held out from training) and a 310-prompt dev stress set. Metrics are measured on the quantized GGUF at temperature 0, without post-hoc regex fixes on model output (raw model JSON only).

Holdout (primary)

Metric Stock Qwen2.5-0.5B Q4_K_M This model (v5)
Task success 50.7% 94.7%
Confidently wrong 34.7% 4.0%
Intent match โ€” 82.7%
Full JSON match (intent + slots) โ€” 64.0%
p50 / p90 latency (Mac eval) โ€” ~499 ms / ~515 ms

Release bar on holdout: task success โ‰ฅ 80%, confidently wrong โ‰ค 8% โ€” met for v5.

Dev (internal, harder)

Task success ~70%, confidently wrong ~26% โ€” includes multi-turn and edge cases; not the primary release gate.

Caveats: Eval scores JSON + routing, not final UX copy. Some merchantโ†’category follow-ups remain imperfect. Quantization may differ slightly from fused F16.


Limitations

  • Domain-specific: personal finance phrasing only.
  • Schema-bound: best results with GBNF or strict JSON validation.
  • No PEFT files: you cannot merge a published LoRA from this repo; only the fused GGUF is available.
  • Training data: synthetic + curated; not a public dataset on the Hub.
  • Not a replacement for licensed financial advice.

License

Apache-2.0, consistent with Qwen2.5-0.5B-Instruct. Use under the same terms as the base model and respect Qwen usage policies.


Citation / links

If you use this model, please cite the base Qwen model and link to this GGUF repository.

Downloads last month
154
GGUF
Model size
0.5B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for pablitoelpeligro/aibudget-intent-05b-GGUF

Quantized
(322)
this model