Instructions to use andyzhang232/ajev-gemma4-12b-lora7 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use andyzhang232/ajev-gemma4-12b-lora7 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
AJev · Gemma 4 12B LoRA (lora7)
AJev is an open decision model in the style of Jev. You give it some context (text or JSON) and a few typed questions: yes / no, single choice (up to 255 options), or an ordered score. It returns a calibrated probability for every option. Each question takes a single forward pass; no text is generated.
This repository holds a LoRA adapter for google/gemma-4-12B-it. The per-type calibration temperatures are in ajev_lm_config.json. Inference code: github.com/cmzy/ajev-infer.
- Needs about 24 GB of GPU memory, so it fits on a single 24–48 GB card.
- If you have enough memory (about 55 GB), the 26B version is recommended: andyzhang232/ajev-gemma4-26b-a4b-lora1, Decision Index 57.42.
Results
Self-run on the full Jev Decision Index 0.2.1 suite with the official kit:
| Model | Base | Decision Index |
|---|---|---|
| AJev 26B-A4B | Gemma 4 26B-A4B | 57.42 |
| AJev lora7 (this model) | Gemma 4 12B | 55.06 |
| AJev lora5 (previous version) | Gemma 4 12B | 52.22 |
| Winnow-12B | Gemma 4 12B | 50.02 |
- By area: Knowledge 38.6, Language 63.4, Retrieval 63.0, Tools 67.5, Arts 37.3.
- These are our own full-suite results (all 150,317 scoreable requests answered, HLE included) and have not yet been reproduced by the maintainers.
- Compared with the previous version, the largest gains are on VAST (stance), HoVer (fact checking), iSarcasm (sarcasm), PhishNChips (phishing emails) and ContractNLI.
Accuracy on held-out test sets:
| Test set | AJev lora5 | AJev lora7 |
|---|---|---|
| typed-decisions | 0.790 | 0.779 |
| JevBench public | 0.853 | 0.844 |
| Kev transfer v9 | 0.771 | 0.769 |
| eikos heldout | 0.931 | 0.927 |
None of the differences between lora7 and lora5 on these test sets are significant.
Usage
pip install "ajev-infer @ git+https://github.com/cmzy/ajev-infer"
from ajev.lm.predictor import LMPredictor
from ajev.schema import decisions_from_jev, jev_answer
p = LMPredictor("google/gemma-4-12B-it", adapter="andyzhang232/ajev-gemma4-12b-lora7")
ds = decisions_from_jev(
{"ticket": "I was charged twice for order #1182."},
{"topic": {"type": "choice", "instructions": "What is the ticket about?",
"criteria": {"billing": "charges, refunds", "delivery": "shipping", "account": "login"}},
"escalate": {"type": "noul", "instructions": "Should a human agent take this now?"}})
for d, probs in zip(ds, p.predict(ds)):
print(d.meta["question_id"], jev_answer(d, probs))
Jev-compatible server (POST /v1/systemone):
pip install "ajev-infer[vllm] @ git+https://github.com/cmzy/ajev-infer"
python -m ajev.serve_vllm --adapter andyzhang232/ajev-gemma4-12b-lora7 --port 8000
- Requires transformers ≥ 5.17.
- Load it as a LoRA adapter, or merge it in memory only. Do not save a merged model and reload it.
Training
- Starting point: the adapter of the previous version, lora5.
- Data: about 36k questions, 43% of them replayed from the previous training data to prevent forgetting. New data includes the training splits of benchmarks related to the leaderboard (HoVer, VAST, POP909, ContractNLI, ACOS, BANKING77, CLINC150) and procedurally generated reasoning questions rewritten into each benchmark's format.
- Decontamination: every training question was checked against the full leaderboard suite, and any with text overlap was removed.
- Setup: LoRA r 32 / α 64, learning rate 1.5e-5, 1 epoch, one RTX PRO 6000.
Limitations
- Tool-use benchmarks (API-Bank, When2Call) are slightly lower than in the previous version.
- Knowledge reasoning (GPQA, ChessBench) and some Arts benchmarks (POP909, cfcolor) are still weak, mainly limited by the 12B base.
- Probabilities were calibrated on our own held-out data. If your data distribution is very different, recalibrate.
- Some training data carries non-commercial or attribution terms. Check them yourself before commercial use.
- Not affiliated with TypeSafe AI.
- Downloads last month
- 17