Qwen3-0.6B-pulse

A 0.6B Qwen3 fine-tuned for one job: being the language layer of Pulse Local, a Mac app where a personal-finance assistant, its ledger and every question stay on the user's own machine. It is published so the app can download it on first launch. Beyond that download, the app's Local Mode makes only an optional check for a newer version, which can be switched off and sends nothing about the user. Its optional Connected Mode adds a bank feed through Plaid; even then the model, the ledger and every question stay on the machine.

Read this before you try it as a general assistant: on its own it is not one. On our private 48-question finance benchmark this model alone scores about 55 out of 100. Inside the app it answers only what code cannot: on that benchmark 41 of the 48 questions are answered in code and 7 reach the model, and on 77 ordinary questions asked of a real ledger the app gives 76 answers with no detectable error, with none of the 77 reaching the model (details and the limits of that scoring below). The tuning is for posture and tool-calling inside that system, not for general capability.

Benchmark: 77 real questions (a test of the app, not of this model)

The questions. 77 ordinary budgeting questions, written the way people ask them in popular budgeting apps, in twelve kinds: spending (14), budget (8), bills (7), cash flow (6), transactions (6), trends (5), time periods (5), multi-step (5), how-to (5), net worth (4), posture (7: "should I cancel Netflix?") and sparse data (5: a category with no transactions, where the only right answer is that there is nothing to report).

The ledger. The founder's own business bank ledger, not synthetic data, so neither the ledger nor the questions are published.

What the scoring checks, and what it does not. No LLM judge. Only 8 of the 77 questions carry expected figures, and for those the answer must contain them. For the other 69 the harness can only catch failures it can see: a figure of $100 or more that no tool, ledger fact or question produced; a month named that is not the one asked about; advice, or a non-answer, where a decision question needs a decline; a total for an empty category; an empty, padded or repeated answer. So the column below counts answers with no detectable error, which is weaker than correct. A wrong answer that states no bad figure passes: gpt-4.1-mini answered "How much did I spend last month?" with "You did not spend any money last month" after querying the wrong year, and the harness passed it.

no detectable error
Pulse Local app with this model (2026-09-27) 76 of 77
the same app with a stock Qwen3-8B (Q4_K_M, 4.99 GB; 2026-09-26) 76 of 77
gpt-4.1-mini with the same tools, no app around it (2026-09-19) 75 of 77
this model alone, with the app's tools but no code in front (2026-09-26) 46 of 77
stock Qwen3-8B alone, same setup (2026-09-26) 52 of 77

How to read it:

  • The first two rows measure the app's code, not the model. In the 2026-09-27 run none of the 77 questions reached the model; every answer came from code. That is why swapping in an 8B changes nothing.
  • The last two rows are the ones about models. Counting an answer with no figure as a miss (the rules above let it pass), they become 35 of 77 for this model and 21 of 77 for the stock 8B. The 8B describes its tool calls instead of making them on 20 of the 77 and asks the user for figures the ledger holds; this model makes the call. That stricter reading was not run on the cloud row.
  • The runs are on different dates, and the harness and the app changed between them, so the rows are indicative, not a controlled comparison.

What this does show: the design keeps the model out of ordinary ledger questions entirely. What it does not show: that this model alone is a good assistant. It is not; see the next section.

How the app keeps it in check

The model is the last step, not the first. Before it is asked anything:

  • ledger questions are planned into a structured query and computed in code;
  • follow-ups ("how did you get that?", "and the average?", "excluding groceries?") are recognised by a small embedding classifier and answered from the previous query in code;
  • a follow-up that nothing recognised is asked again as a fresh question, or answered with a written menu, never by the model.

And after it answers:

  • any money figure in its answer must come from the question, the conversation, the ledger or a tool result, otherwise the answer is replaced with a written one;
  • an answer that names a financial product the question did not name is replaced with a written one.

The embedding model used for routing and follow-ups is bge-small-en-v1.5 (GGUF from CompendiumLabs), downloaded alongside this one.

What the tuning changed

base Qwen3-0.6B this model
posture (describes, never directs) 30 90
answers that told the user to buy or sell 2 0
tool calls in the app's format kept kept
answer length long about 4x shorter

"Posture" is the rule the app exists to keep: Pulse describes a financial situation and never tells anyone to buy, sell, hold, trim or allocate. The base model failed it, answering "selling now could be a good move" to a question about a falling holding. That is the failure this tuning exists to fix.

Files

file size use
Qwen3-0.6B-pulse-v4-Q4_K_M.gguf 397 MB what the app ships; about 120 tokens/s on an M4
Qwen3-0.6B-pulse-v4-Q8_0.gguf 639 MB slightly better quality, slower

Q6_K was measured and is not better than Q4_K_M in the app, so it is not published.

How it is meant to be run

  • Chat template: Qwen3's, with the empty thinking block not prefilled. This model was trained on a template that emits no <think> block; prefill one and it drifts.
  • Tool calls: a plain generation loop that parses <tool_call>{"name": ..., "arguments": {...}}</tool_call> from the text. Note that node-llama-cpp's built-in function calling drops this model's calls; the same GGUF works in raw llama.cpp and MLX. The app uses its own loop for this reason.
  • Sampling: greedy (temperature 0). Sampled, it derails mid-JSON.
  • Context: 4096 is plenty; the app never needs more.

Training

  • Base: Qwen/Qwen3-0.6B, Apache 2.0.
  • Method: LoRA (rank 16, 16 layers, lr 3e-5) with mlx-lm on an M4 Mac, then fused and converted to GGUF with llama.cpp.
  • Data: about 1,860 synthetic rows, roughly 1,700 written answers plus tool-call rows generated from the app's own tool schemas and calculator. The answers were distilled from gpt-4.1-mini and filtered for directive phrasing, boilerplate, em dashes and length. No personal or user data was involved at any point; every ledger figure in the training data is synthetic.
  • A later round on 5,617 rows was trained and measured. It was flat, 8 items better and 9 worse out of 348, so it is not published.

Limitations

  • English only, US personal finance only.
  • No internet and no live data. It is trained to say so rather than invent a price, a rate or an index level.
  • It is bad at one-word yes/no answers: measured 13 correct out of 30 on questions with an unambiguous answer, biased toward "yes". The app never asks it for one.
  • It gets arithmetic method right and numbers wrong, which is why the app computes every figure in code rather than trusting the model.
  • It is not a financial adviser and neither is the app.

Where it runs

The app that carries this model is downloadable at pulseos.finance (Mac, Apple Silicon, signed and notarised, beta). It is pay what you want, including nothing, and the weights here are open under the license below. The model has not changed since 2026-09-19; the app around it has (a bank feed, a query planner, a follow-up classifier, the number check above, a Japanese interface), which is why the app's score keeps moving while the model's does not. Card updated 2026-09-27.

Provenance note

The training answers were generated with OpenAI models. If you intend to build on this, check OpenAI's terms for your own use case.

License

Apache 2.0, inherited from the base model, whose LICENSE is included.

Downloads last month
120
GGUF
Model size
0.6B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for PulseOS/pulse-local-qwen3-0.6b

Finetuned
Qwen/Qwen3-0.6B
Quantized
(460)
this model