decosa-mt-term-lora-eurollm-9b

A LoRA adapter for EuroLLM-9B-Instruct-2512 that makes glossary-constrained translation keep its grammar. When a translation must use a given term ("recall translates to terugroepen"), a model tends to paste the dictionary form into the sentence, even where the sentence needs another case, number, article or agreement. This adapter was trained to use the given term in the form the sentence needs.

English into German, French, Spanish, Italian, Dutch and Polish. Use it only for prompts that carry terms; send prompts without terms to the base model (see Limitations: it is slightly worse than the base on general text).

Results

Measured on 28 Sep 2026 on one vLLM 0.29.0 server with the FP8 base model and this adapter loaded with --lora-modules, so base and adapter share everything else. Greedy decoding, repetition penalty 1.0. Hy-MT2-7B rows are Tencent's translation model on the same test, with and without our own term adapter for it.

Held-out terms in held-out documents (the main test). 644 sentences of EU legislation from DGT-TM (documents never used in training) with 658 IATE terms whose concepts were never a training constraint. "Exact" means the output contains the professional translator's inflected form of the term; "present" means any form of the term is there.

Model Term present Exact inflected form Dictionary form pasted where the sentence needs another
EuroLLM-9B-Instruct (base) 618 (93.9%) 535 (81.3%) 59
EuroLLM-9B + this adapter 639 (97.1%) 580 (88.1%) 35
Hy-MT2-7B (base) 641 (97.4%) 528 (80.2%) 82
Hy-MT2-7B + our Hy-MT2 term adapter 630 (95.7%) 549 (83.4%) 57
  • Paired per term against the base: 55 terms gained, 10 lost for "exact" (two-sided exact binomial p = 1e-8); 26 gained, 5 lost for "present".
  • Exact by language (base -> adapter): de 105 -> 110 of 138, fr 105 -> 112 of 122, es 65 -> 76 of 83, it 83 -> 93 of 106, nl 48 -> 55 of 61, pl 129 -> 134 of 148.
  • The same test on the training machine (bf16, adapter merged, transformers) gave 536 -> 577 exact: FP8 serving does not change the result.
  • This test is in the training domain (EU legislation), so it shows what the adapter learned under the best conditions.

Out of domain: clinical-trial lay text and product-safety text with a glossary forcing 435 terms (40 sentences x 6 languages, one call, no retries, graded by the glossary's rules):

Model Terms used correctly Missing Forbidden form Sentences with every term right
EuroLLM-9B (base) 423 / 435 9 3 229 / 240
+ this adapter 427 / 435 6 2 233 / 240

We read all 198 sentence pairs that differ (not native speakers). The adapter fixed 10 errors of the base (for example "vehicle" rendered as a car instead of the forced "placebo", a German sentence that swapped "recall" and "withdrawal", an Italian sentence that reversed who informs whom) and introduced 6 (for example Polish "importer" translated as "user", German "Dieses kurze Zusammenfassung"). A small net gain here, not the large one of the main test.

General text, FLORES-200 devtest (1,012 sentences per language, every prompt sent through the adapter, no terms):

chrF++ / COMET-22 de fr es it nl pl
EuroLLM-9B (base) 63.73 / 88.78 68.22 / 88.91 54.76 / 87.32 57.83 / 89.37 55.56 / 88.75 50.41 / 90.49
+ this adapter 61.96 / 88.35 67.55 / 88.76 53.98 / 87.11 56.88 / 89.20 55.22 / 88.66 50.03 / 90.22

The adapter is worse on text without terms (-0.3 to -1.8 chrF++, -0.1 to -0.4 COMET-22). That is why it should only see term prompts.

All numbers are in eval_summary.json.

Intended use

  • Glossary- or term-bank-constrained machine translation from English into de, fr, es, it, nl, pl, with EuroLLM-9B, where the terms must appear and the sentence must stay grammatical (regulatory, clinical, product-safety text).
  • As one step in a pipeline that still checks the output: that every forced term appears, that numbers and negations are kept, and a human review where the text matters. The adapter reduces errors; it does not remove them.

Not for: translation without terms (use the base model), other language pairs (untested), or any clinical, legal or safety text published without review.

Limitations (measured, see Results)

  • Worse on general text than the base when used without terms, and it drifts towards the wording of EU legislation ("Arzneimittel" for "medicine", "Per «evento avverso» si intende ..."). Correct, but more formal than lay text wants.
  • Trained on one domain (EU legislation from 2020). The gain out of domain is small (427 vs 423 of 435) and it adds a few errors of its own; the large gain is in-domain.
  • It sometimes drops a term the base keeps (10 of 658 "exact" and 5 "present" on the main test; Dutch "present" 56 -> 55). Check for the term and retry, or fall back to the base, when it is missing.
  • German and Dutch still misread "vehicle" (as in placebo vehicle) as a car; Polish sometimes leaves an English term ("Recall") untranslated.
  • No native-speaker review of the outputs.

Training

Base utter-project/EuroLLM-9B-Instruct-2512 @ def82454 (Apache-2.0), bf16
Method LoRA rank 16, alpha 32, dropout 0.05 on q/k/v/o/gate/up/down in all 42 layers (50.9 M parameters, stored fp32); loss on the answer only
Schedule 2 epochs, 1,725 steps, batch 16, max length 640, AdamW, LR 1e-4 cosine with 50 warmup steps; one H200, 26 min
Examples 13,800 (417 validation): per language 1,800 term prompts with 1 to 3 terms (at least 75% need an inflected form), 350 plain prompts and 150 repair prompts; validation loss 0.610 -> 0.391
Held out 4% of IATE concepts (never a training constraint) and 4% of DGT-TM documents (never in training); the main test is built from both

Training data (every source and its licence)

Source What was used Licence Attribution
DGT-TM 2021, volume Vol_2020_1 English sentences of EU acts adopted in 2020 and their human translations (targets) Commission Decision 2011/833/EU: reuse for commercial or non-commercial purposes with acknowledgement © European Union, European Commission DG Translation / JRC
IATE (full TBX export IATE_export_26022019) English terms and their de/fr/es/it/nl/pl equivalents (reliability 3+, not deprecated), matched to the DGT-TM sentences Commission Decision 2011/833/EU, as listed on data.europa.eu © European Union, EU institutions (IATE)

Every target is a professional human translation from DGT-TM; no model-written text was used for training. The training and evaluation data are not redistributed here.

Usage

pip install torch transformers peft accelerate
python usage.py

The prompt the adapter was trained on (one X translates to Y line per term):

Translate the following English source text to Dutch. Use exactly these term translations, adapting only their grammatical form:
market surveillance authority translates to markttoezichtautoriteit
recall translates to terugroepen

English: The market surveillance authority ordered a recall of the toy because of a choking risk. 
Dutch: 

Sent as the user message of EuroLLM's chat template. With vLLM: --enable-lora --max-lora-rank 16 --lora-modules eurollm-9b-term=decosaai/decosa-mt-term-lora-eurollm-9b (local path or Hub id), then request model eurollm-9b-term for term prompts and the base model for everything else.

Files: adapter_model.safetensors, adapter_config.json (PEFT), usage.py, eval_summary.json, NOTICE, LICENSE, SHA256SUMS.

Licence

Apache-2.0 for the adapter weights and code in this repository. The base model is Apache-2.0 and is not included. The training data sources above are credited here and in NOTICE.

Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for decosaai/decosa-mt-term-lora-eurollm-9b