Instructions to use pablitoelpeligro/aibudget-intent-05b-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use pablitoelpeligro/aibudget-intent-05b-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Use Docker
docker model run hf.co/pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use pablitoelpeligro/aibudget-intent-05b-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pablitoelpeligro/aibudget-intent-05b-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pablitoelpeligro/aibudget-intent-05b-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
- Ollama
How to use pablitoelpeligro/aibudget-intent-05b-GGUF with Ollama:
ollama run hf.co/pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use pablitoelpeligro/aibudget-intent-05b-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use pablitoelpeligro/aibudget-intent-05b-GGUF with Docker Model Runner:
docker model run hf.co/pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
- Lemonade
How to use pablitoelpeligro/aibudget-intent-05b-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.aibudget-intent-05b-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use pablitoelpeligro/aibudget-intent-05b-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use pablitoelpeligro/aibudget-intent-05b-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "pablitoelpeligro/aibudget-intent-05b-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LoRA fine-tune of Qwen2.5 0.5B (~398 MB Q4_K_M) for phones: small enough to download over cellular and run on-device, specialized for personal-finance intent JSON (EN/IT/ES/FR).
AiBudget Intent 0.5B v5
On-device personal-finance intent โ strict JSON (LoRA SFT on Qwen2.5-0.5B-Instruct, fused, Q4_K_M GGUF)
Maps natural-language questions about personal spending (EN / IT / ES / FR) into a single, schema-constrained JSON intent object.
The AiBudget mobile app parses the JSON and runs local SQLite tools. When users choose this on-device model, transaction data is not sent to a cloud LLM for intent parsing.
Why a 0.5B model? We deliberately started from a very small instruct model so the shipped artifact stays around one download (~400 MB) and can run locally on mid-range phones without a cloud API. LoRA SFT narrows that base model to one jobโstructured budget intentsโnot open-ended chat. Larger models can be smarter in general, but they are harder to justify for privacy-first, offline-friendly mobile installs.
Published artifact: fused, quantized GGUF only โ LoRA adapter weights are not in this repository.
Recommended download: aibudget-intent-05b-v5-Q4_K_M.gguf (~398 MB, SHA256 d3d6fc0780359cd4d198e9354b9ecfc0ff9cfb0fb729f040326fa6e2f02c693a)
Other aibudget-intent-05b-v*-Q4_K_M.gguf files in this repo are experimental rounds; use v5 unless you are reproducing eval history.
Model details
| Developed by | pablitoelpeligro (AiBudget) |
| Model type | Causal LM; LoRA SFT โ merged weights โ GGUF Q4_K_M |
| Languages | English, Italian, Spanish, French |
| License | Apache-2.0 (same as base) |
| Base model | Qwen/Qwen2.5-0.5B-Instruct |
| Training | mlx-lm on Apple Silicon โ fuse โ llama.cpp Q4_K_M |
| Production inference | llama.rn (React Native) + GBNF grammar matching the intent JSON schema |
Intended use
Direct use
Drop-in intent parser for personal-finance assistants (especially mobile / edge):
- Total spend, category spend, merchant (
description) spend - Transaction search, sorting, limits
- Budget status, monthly trends, top merchants
- Period comparisons (
compare_periods) - Structured
add_transactionextraction unknownfor meta / out-of-scope questions
Trained and evaluated with the same compact prompt profile as production (short system text, no few-shot examples in the prompt).
Out-of-scope
- General open-ended chat
- Nonโpersonal-finance domains
- Financial advice without human review
- Languages outside EN / IT / ES / FR
How to use
Download
hf download pablitoelpeligro/aibudget-intent-05b-GGUF \
aibudget-intent-05b-v5-Q4_K_M.gguf --local-dir .
Prompt format (compact)
Production builds chat messages like this (Qwen chat template):
- System: compact classifier text + newline +
conversationContext: {"today":"โฆ",โฆ}(JSON) - Optional: recent user turns only (no assistant prose in history)
- User: current question
Example system line (English app, English query):
Translate the user's message into exactly one AiBudget financial intent. Use conversationContext and earlier user messages to resolve follow-ups. Message language: en. App language: en. Return only JSON.
conversationContext: {"today":"2026-09-22","currency":"USD"}
Use temperature: 0 for reproducible decoding (matches our offline eval).
llama.cpp (best-effort without grammar)
Without GBNF, the model may still emit JSON but slot accuracy can drop vs the app. For a quick smoke test:
# Flag names vary by llama.cpp version; use Qwen2.5-Instruct chat template in production.
llama-cli -m aibudget-intent-05b-v5-Q4_K_M.gguf --temp 0 -n 512 \
-sys "Translate the user's message into exactly one AiBudget financial intent. Return only JSON." \
-p "Total spending this month"
AiBudget app
In Settings, download AiBudget Intent 0.5B (tuned v5). The app loads this GGUF with llama.rn and grammar-constrained decoding so output matches the schema below.
Output schema
One JSON object per turn. All keys required; unused slots are null. intent and dateRange are always strings (never null). confidence is in [0, 1].
| Field | Type | Role |
|---|---|---|
intent |
enum | total_spend ยท description_spend ยท category_spend ยท category_breakdown ยท monthly_trends ยท budget_status ยท transaction_search ยท add_transaction ยท net_cashflow ยท compare_periods ยท top_descriptions ยท unknown |
description |
string | null | Merchant / payee (1โ4 words) |
category |
string | null | Category label |
dateRange |
enum | this_month ยท last_month ยท today ยท yesterday ยท previous_range ยท last_n_months ยท this_year ยท last_year ยท year_to_date ยท quarter ยท custom ยท all_time ยท unspecified |
monthsCount |
int | null | For last_n_months |
startDate, endDate |
string | null | Custom range when dateRange is custom |
periodLabel |
string | null | Free-text period hint |
amount |
number | null | e.g. search by amount |
notes |
string | null | e.g. add_transaction |
sort |
enum | null | amount_desc ยท amount_asc ยท date_desc ยท date_asc |
limit |
int | null | Row cap for lists / tops |
metric |
enum | null | sum ยท max ยท min ยท count |
transactionType |
enum | null | expense ยท income |
ordinal |
int | null | 1 = most recent in sort order |
offset |
int | null | Pagination offset |
quarterNumber |
int | null | 1โ4 when dateRange is quarter |
year |
int | null | Year for quarter / year scopes |
dateRangeA, dateRangeB |
enum | null | Same values as dateRange; period comparison |
startDateB, endDateB, periodLabelB |
string | null | Second period for custom compare |
groupBy |
enum | null | category ยท description |
compareMetric |
enum | null | sum ยท delta ยท pct_change |
confidence |
number | Model confidence |
Example
User: how much did i blow on Uber last month
{
"intent": "description_spend",
"description": "Uber",
"category": null,
"dateRange": "last_month",
"monthsCount": null,
"startDate": null,
"endDate": null,
"periodLabel": null,
"amount": null,
"notes": null,
"sort": null,
"limit": null,
"metric": null,
"transactionType": null,
"ordinal": null,
"offset": null,
"quarterNumber": null,
"year": null,
"dateRangeA": null,
"dateRangeB": null,
"startDateB": null,
"endDateB": null,
"periodLabelB": null,
"groupBy": null,
"compareMetric": null,
"confidence": 0.9
}
Training
| Setting | Value |
|---|---|
| Framework | mlx-lm (mlx_lm.lora) |
| Base | Qwen/Qwen2.5-0.5B-Instruct (full precision, not QLoRA) |
| Method | LoRA rank 8, 16 layers, dropout 0.05, MLX scale 20, lr 2e-5 |
| Steps | 1200 iterations, batch 4 ร grad accum 2, max seq 1536 |
| Loss | Response only (mask_prompt: true) |
| Export | Fuse adapter โ GGUF F16 โ Q4_K_M (evaluated artifact) |
| Hardware | Apple Silicon (MLX) |
Data (not on Hugging Face): ~800 hand-written seed rows (EN/IT/ES/FR), ~83 curated examples, ~2k synthetic LLM-generated rows (schema- and leak-checked), plus small targeted rounds (v5 added ~115 rows for unknown / category narrows). Typical MLX split: ~2.5kโ2.8k train / ~280โ310 valid. Training chat format matches production (compact system prompt, no in-prompt few-shots).
Evaluation
Offline eval on a 75-prompt holdout (held out from training) and a 310-prompt dev stress set. Metrics are measured on the quantized GGUF at temperature 0, without post-hoc regex fixes on model output (raw model JSON only).
Holdout (primary)
| Metric | Stock Qwen2.5-0.5B Q4_K_M | This model (v5) |
|---|---|---|
| Task success | 50.7% | 94.7% |
| Confidently wrong | 34.7% | 4.0% |
| Intent match | โ | 82.7% |
| Full JSON match (intent + slots) | โ | 64.0% |
| p50 / p90 latency (Mac eval) | โ | ~499 ms / ~515 ms |
Release bar on holdout: task success โฅ 80%, confidently wrong โค 8% โ met for v5.
Dev (internal, harder)
Task success ~70%, confidently wrong ~26% โ includes multi-turn and edge cases; not the primary release gate.
Caveats: Eval scores JSON + routing, not final UX copy. Some merchantโcategory follow-ups remain imperfect. Quantization may differ slightly from fused F16.
Limitations
- Domain-specific: personal finance phrasing only.
- Schema-bound: best results with GBNF or strict JSON validation.
- No PEFT files: you cannot merge a published LoRA from this repo; only the fused GGUF is available.
- Training data: synthetic + curated; not a public dataset on the Hub.
- Not a replacement for licensed financial advice.
License
Apache-2.0, consistent with Qwen2.5-0.5B-Instruct. Use under the same terms as the base model and respect Qwen usage policies.
Citation / links
- This repository: https://huggingface.co/pablitoelpeligro/aibudget-intent-05b-GGUF
- Base model: https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct
If you use this model, please cite the base Qwen model and link to this GGUF repository.
- Downloads last month
- 154
4-bit