Instructions to use teex-pt/AMALIA-9B-0626-DPO-LoRA-honesty-pilot with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use teex-pt/AMALIA-9B-0626-DPO-LoRA-honesty-pilot with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("teex-pt/AMALIA-9B-0626-DPO-LoRA-honesty-pilot") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use teex-pt/AMALIA-9B-0626-DPO-LoRA-honesty-pilot with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "teex-pt/AMALIA-9B-0626-DPO-LoRA-honesty-pilot" --prompt "Once upon a time"
- Atomic Chat
AMALIA-9B LoRA — merge-75 (pilot series champion)
LoRA adapter from a verifier-gated fine-tuning pilot on AMALIA-9B-0626-DPO, targeting confabulation (invented biographies, fake-entity answers, future-event "knowledge") while protecting arithmetic and instruction-following. Full methodology, harness, and every iteration's report: github.com/teex-pt/pt-amalia.
Update (2026-07-07): this repo now ships merge-75, found by mapping the
merge's α curve more finely — it supersedes both the original v2 adapter
and the earlier merge-65v2 published here. v2's dataset is still
available at teex-pt/amalia-pilot-honesty-v2
for anyone who wants to reproduce the original run.
What it is: a weighted merge, not a new training run
merge-75 is a weighted average of two LoRA adapters' weights (α=0.75
toward v2, 0.25 toward v4) — pure vector arithmetic, zero additional
training or generation. adapter_config.json's training metadata (iters,
data, etc.) reflects v2's original run since the merge script uses it as a
template; the actual weights are the blend, not from that single run.
The pilot series (what led here)
| Iteration | What changed | Verdict |
|---|---|---|
| v1 | Refusal-only SFT (440 samples) | Rejected — +43pp honesty but arithmetic collapsed (−23pp) and over-refused real entities |
| v2 | Mixed data: refusals + on-policy real-QA + verified anchors | Accepted — honesty 96%, arithmetic 49%, control 100% |
| v3 | v2 recipe + reasoning-style arithmetic anchors | Rejected as overall winner, but proved anchor style transfers to task style (GSM8K CoT +16pp, series-best IFEval 68%) |
| v4 | Both anchor styles (bare + reasoned), matched to instruction | Best arithmetic (52%) and best GSM8K CoT (66%) of the series |
| merge-65v2 | Weighted average of v2 × v4, α=0.65 | Best all-round harness score at the time — but gave up most of v4's CoT gain |
| merge-75 | Same v2 × v4 pair, α=0.75 — finer sweep of the blend ratio | This repo — dominates v2 on every harness axis, ties v3's series-best IFEval |
Full results (extended harness, n=100/30/36, plus consortium CoT tasks)
| Metric | Baseline | v2 | v3 | v4 | merge-65v2 | merge-75 |
|---|---|---|---|---|---|---|
| honesty | 50.0% | 96.0% | 81.0% | 82.0% | 94.0% | 96.0% (ties v2) |
| arithmetic (answer-only) | 46.0% | 49.0% | 36.0% | 52.0% | 50.0% | 51.0% |
| format | 73.3% | 80.0% | 73.3% | 73.3% | 76.7% | 80.0% (ties v2) |
| variety | 86.7% | 93.3% | 93.3% | 86.7% | 90.0% | 93.3% (ties v2/v3) |
| control (36 real entities) | 100% | 100% | 97.2% | 100% | 100% | 100% |
| overall | 55.4% | 75.8% | — | 70.0% | 74.6% | 76.5% — best of series |
| GSM8K-pt CoT | 48.0%¹ | 48.0%¹ | 64.0% | 66.0% | 52.0% | 54.0% |
| IFEval-pt strict | 60.0% | — | 68.0% | 64.0% | 64.0% | 68.0% (ties v3) |
¹ measured at n=25 for baseline/v2, n=50 for v3/v4/merge-65v2/merge-75.
merge-75 dominates v2 on every axis measured — matches its honesty, format, and variety exactly, while beating its arithmetic (51.0% vs 49.0%) and overall harness score (76.5% vs 75.8%). It also ties v3's series-best IFEval score. It doesn't lead on raw arithmetic (v4 is 1pp higher, within noise at n=100) or GSM8K-CoT (v3/v4's reasoning-anchor training still leads there) — pick v4 specifically if your use case is reasoning/CoT-heavy; otherwise this is the strongest general-purpose checkpoint of the series.
How it was made
- v1→v2: templated refusals for fabricated entities/future events + on-policy real-entity QA (verified, not just assumed correct) + verified arithmetic/format anchors from a Ministral-3-14B-Reasoning teacher.
- v3→v4: same recipe, with arithmetic anchors in two styles (bare answer vs short-reasoning), each paired with a matching instruction.
- merge:
scripts/merge_adapters.py—merged[k] = α·v2[k] + (1-α)·v4[k]for every LoRA tensor. Swept α∈{0.50, 0.55, 0.60, 0.65, 0.70, 0.75} on a subset; 0.75 won on both arithmetic and honesty simultaneously.
Every dataset sample behind every step is gated by deterministic code verifiers (arithmetic ground truth computed by templates, not models; honesty targets checked for uncertainty markers and confabulation patterns) — no LLM judging anywhere in the data pipeline.
Usage
pip install mlx-lm
mlx_lm.generate --model amalia-llm/AMALIA-9B-0626-DPO --adapter-path <this-repo> \
--prompt "Quem foi o poeta Aurélio Vasconcelos de Mirandela?"
# base model: invents a biography; with this adapter: honestly says it doesn't know
Attribution
Base model by the AMALIA team (Apache 2.0). Adapter, data, and method: teex-pt.
Quantized
Model tree for teex-pt/AMALIA-9B-0626-DPO-LoRA-honesty-pilot
Base model
amalia-llm/AMALIA-9B-0626-SFT