🚀 Official Release: FiscMind-Qwen38-27B-RLVR - Verifiable Reward Alignment (GRPO) for Romanian Accounting & Fiscal Compliance
🎯 Announcement: FiscMind-Qwen38-27B-RLVR Official Community Release
We are proud to present FiscMind-Qwen38-27B-RLVR, the reinforcement-aligned flagship model for Romanian fiscal jurisprudence and SAGA C double-entry accounting automation.
🏛️ Why Reinforcement Learning with Verifiable Rewards (RLVR)?
Traditional Supervised Fine-Tuning (SFT) often struggles with the strict mathematical and statutory constraints of Romanian tax law (Legea 227/2015 Codul Fiscal) and accounting standards (OMFP 1802/2014):
- Deterministic Accounting Balances: In accounting, sum(Debit) = sum(Credit) cannot tolerate approximations (Delta = 0.00 RON must be exact).
- Statutory Guardrails: Deductibility limits (e.g. 50% mixed-use auto under Art. 68/Art. 298, 5% gross profit for legal reserves up to 20% capital under Art. 26) must cite the exact legal articles.
- Safety & Human Oversight: In accordance with the project safety constitution (AGENTS.md), the model must refuse automatic final approval, requiring human confirmation before executing or exporting journal entries.
To solve this, FiscMind-Qwen38-27B-RLVR was trained using Group Relative Policy Optimization (GRPO) with 4 deterministic verifiable reward functions:
- R_SMT (Double-Entry Balance & OMFP 1802 Accounts) [Weight: 0.35]: Verified mathematically using SMT balance constraints.
- R_Law (Statutory Citations) [Weight: 0.25]: Strict validation of fiscal code and government ordinance citations.
- R_Format (Deliberative CoT ...) [Weight: 0.25]: Structured thought process before emitting the accounting journal.
- R_Safety (AGENTS.md Strict Guardrails) [Weight: 0.15]: Absolute prevention of unauthorized accounting execution.
📊 Benchmark Results: Baseline vs. Post-RLVR
Across all 8 core Romanian fiscal archetypes (AIC, Auto mixt 50%, Stat salarii, Închidere TVA, Amortizare liniară, Tichete de masă OUG 115/2023, Rezervă legală, Vânzare & descărcare mărfuri):
| Metric | Baseline CoT SFT | FiscMind RLVR (Policy) | Improvement |
|---|---|---|---|
| R_SMT (Mathematical Balance Delta = 0.00) | 50.0% | 100.0% | +50.0% |
| R_Law (Cod Fiscal & OMFP 1802 Citation) | 70.0% | 87.5% | +17.5% |
| R_Format (Structured Deliberative CoT) | 40.6% | 100.0% | +59.4% |
| R_Safety (Human Approval & Refusal Guardrails) | 100.0% | 100.0% | 100.0% (Maintained) |
| COMPOSITE RLVR SCORE | 60.2% | 96.8% | +36.6% |
⚡ Quickstart with Transformers & PEFT
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
BASE_MODEL = "Qwen/Qwen3.8-27B"
ADAPTER_REPO = "iulio/FiscMind-Qwen38-27B-RLVR"
tokenizer = AutoTokenizer.from_pretrained(ADAPTER_REPO)
base_model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL,
torch_dtype=torch.bfloat16,
device_map="auto"
)
model = PeftModel.from_pretrained(base_model, ADAPTER_REPO)
messages = [
{"role": "system", "content": "Ești FiscMind AI, expert contabil autorizat CECCAR conform OMFP 1802/2014 și Codul Fiscal."},
{"role": "user", "content": "Societatea achiziționează combustibil 500 lei + TVA 19% pentru autoturism utilizat mixt (50%). Prezintă monografia și deducerea fiscală."}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.6)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
💼 SAGA C Integration & Local Architecture
The model natively formats accounting entries into SAGA C XML (<OperatiuniDiverse>) and pipe-delimited text import formats, backed by local unit tests (296/296 passing).
Feedback, real-world case submissions, and discussions are welcome!