🚀 Official Release: FiscMind-Qwen38-27B-RLVR - Verifiable Reward Alignment (GRPO) for Romanian Accounting & Fiscal Compliance

#1
by iulio - opened

🎯 Announcement: FiscMind-Qwen38-27B-RLVR Official Community Release

We are proud to present FiscMind-Qwen38-27B-RLVR, the reinforcement-aligned flagship model for Romanian fiscal jurisprudence and SAGA C double-entry accounting automation.


🏛️ Why Reinforcement Learning with Verifiable Rewards (RLVR)?

Traditional Supervised Fine-Tuning (SFT) often struggles with the strict mathematical and statutory constraints of Romanian tax law (Legea 227/2015 Codul Fiscal) and accounting standards (OMFP 1802/2014):

  • Deterministic Accounting Balances: In accounting, sum(Debit) = sum(Credit) cannot tolerate approximations (Delta = 0.00 RON must be exact).
  • Statutory Guardrails: Deductibility limits (e.g. 50% mixed-use auto under Art. 68/Art. 298, 5% gross profit for legal reserves up to 20% capital under Art. 26) must cite the exact legal articles.
  • Safety & Human Oversight: In accordance with the project safety constitution (AGENTS.md), the model must refuse automatic final approval, requiring human confirmation before executing or exporting journal entries.

To solve this, FiscMind-Qwen38-27B-RLVR was trained using Group Relative Policy Optimization (GRPO) with 4 deterministic verifiable reward functions:

  1. R_SMT (Double-Entry Balance & OMFP 1802 Accounts) [Weight: 0.35]: Verified mathematically using SMT balance constraints.
  2. R_Law (Statutory Citations) [Weight: 0.25]: Strict validation of fiscal code and government ordinance citations.
  3. R_Format (Deliberative CoT ...) [Weight: 0.25]: Structured thought process before emitting the accounting journal.
  4. R_Safety (AGENTS.md Strict Guardrails) [Weight: 0.15]: Absolute prevention of unauthorized accounting execution.

📊 Benchmark Results: Baseline vs. Post-RLVR

Across all 8 core Romanian fiscal archetypes (AIC, Auto mixt 50%, Stat salarii, Închidere TVA, Amortizare liniară, Tichete de masă OUG 115/2023, Rezervă legală, Vânzare & descărcare mărfuri):

Metric Baseline CoT SFT FiscMind RLVR (Policy) Improvement
R_SMT (Mathematical Balance Delta = 0.00) 50.0% 100.0% +50.0%
R_Law (Cod Fiscal & OMFP 1802 Citation) 70.0% 87.5% +17.5%
R_Format (Structured Deliberative CoT) 40.6% 100.0% +59.4%
R_Safety (Human Approval & Refusal Guardrails) 100.0% 100.0% 100.0% (Maintained)
COMPOSITE RLVR SCORE 60.2% 96.8% +36.6%

⚡ Quickstart with Transformers & PEFT

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

BASE_MODEL = "Qwen/Qwen3.8-27B"
ADAPTER_REPO = "iulio/FiscMind-Qwen38-27B-RLVR"

tokenizer = AutoTokenizer.from_pretrained(ADAPTER_REPO)
base_model = AutoModelForCausalLM.from_pretrained(
    BASE_MODEL,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)
model = PeftModel.from_pretrained(base_model, ADAPTER_REPO)

messages = [
    {"role": "system", "content": "Ești FiscMind AI, expert contabil autorizat CECCAR conform OMFP 1802/2014 și Codul Fiscal."},
    {"role": "user", "content": "Societatea achiziționează combustibil 500 lei + TVA 19% pentru autoturism utilizat mixt (50%). Prezintă monografia și deducerea fiscală."}
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.6)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

💼 SAGA C Integration & Local Architecture

The model natively formats accounting entries into SAGA C XML (<OperatiuniDiverse>) and pipe-delimited text import formats, backed by local unit tests (296/296 passing).

Feedback, real-world case submissions, and discussions are welcome!

Sign up or log in to comment