HinoMoto-350M Cultural-SFT v1

HinoMoto-350M (from-scratch Japanese LM) を HinoMoto-Bench-ja v0.6 で Cultural SFT した merged 版.

「家族・敬語・沈黙・思いやり・地方文化・ニュアンス・世代差・職場」 の 8 文化軸での応答に特化.

What this is

  • base: HinoMoto-350M v1 (318M params, from-scratch 日本語 LM, vocab 9506)
  • fine-tune: LoRA r=32 / alpha=64 / 5 epoch
  • data: HinoMoto-Bench-ja v0.6 train split (551 items / 8 cultural axes)
  • adapter merged into base (single safetensors file, no peft loading needed)

SFT 効果 (40 items eval, 8軸×5)

Metric base alone +Cultural SFT (this model)
gold bigram overlap 0.039 0.181 (+364%)
mean output length 104 chars 102 chars (適切に保持)
artifact rate 14/40 (35%) 1/40 (2.5%)
mean_token_accuracy (train end) — 86.1%
train_loss (final) — 0.78

→ from-scratch 350M base に対して gold overlap が 3.6x 改善 + artifact ほぼ除去.

Training summary

Item Value
Method LoRA (r=32, alpha=64, target=all linears)
Epochs 5
Batch 2 × grad_accum 4 (effective 8)
Seq len 512
LR 3e-4, cosine schedule, warmup 10%
Precision bf16 mixed
Hardware RTX 3090 24GB
Wall-clock ~10 min
Data 551 train items (80% of bench v0.6)
Eval split 141 held-out items

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "FiShota/hinomoto-350m-cultural-sft-v1",
    dtype=torch.bfloat16,
)
tok = AutoTokenizer.from_pretrained("FiShota/hinomoto-350m-cultural-sft-v1")

prompt_template = """{axis}軸: {axis_def}
場面: {context}
ユーザー発言: {user}
回答:"""

# Example: keigo axis
prompt = prompt_template.format(
    axis="keigo",
    axis_def="敬語 (尊敬・謙譲・丁寧) の使い分けが問われる場面",
    context="上司に報告メールを書く場面",
    user="昨日の会議について",
)

inputs = tok(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=80, do_sample=False, top_k=1)
print(tok.decode(out[0], skip_special_tokens=True))

License

CC BY 4.0. Use for any purpose with attribution.

Intended Uses & Limitations

Intended uses:

  • Research on Japanese cultural-axis evaluation
  • Demonstration of from-scratch JP LM + LoRA SFT on consumer GPU
  • Educational reference for small-LM bench-trained models

Out-of-scope:

  • Production assistant / chatbot (no safety alignment, limited corpus)
  • Tasks requiring broad world knowledge (78MB pre-training corpus is small)
  • Code generation, math, factual QA outside the cultural-axis domain

Bias, Risks, Limitations

  • n=1 run (seed=0 only) — variance not measured
  • Bench-trained: works best on cultural-axis prompts in bench v0.6 format. For general chat, may be suboptimal.
  • No safety alignment. May produce inappropriate content if prompted out-of-distribution.
  • Limited corpus (350M base trained on 78MB; SFT on 551 items): expect mistakes & repetitions on out-of-distribution prompts.
  • Japanese cultural perspective embedded: bench items reflect a specific (urban Japanese) view; minority dialects / regional variations underrepresented.

Transparency: Bench / Corpus Leak Audit

The training corpus and HinoMoto-Bench-ja have been audited for n-gram overlap (8-char):

Bench axis mean n-gram leak max items >30% leak
family 2.7% 51.8% 2 / 110
keigo 22.9% 77.8% 25 / 70 (※ note below)
silence 1.4% 27.3% 0 / 50

Note on keigo: high overlap is structural — top "leaks" are standard formulaic Japanese expressions like 「今しばらくお待ちください」「おはようございます」 that appear naturally in any Japanese corpus and are exactly the target output of the keigo evaluation axis. This is not contamination; it reflects the language's reliance on standardized polite forms.

The bench/corpus pair is considered publication-ready under this analysis.

Audit script: scripts/bench_leak_audit.py in the training repo.

Related

Citation

@misc{hinomoto-350m-cultural-sft-2026,
  author = {ryu (FIshota)},
  title = {HinoMoto-350M Cultural-SFT v1},
  year = {2026},
  publisher = {Hugging Face},
}

Generated: 2026-05-13 (HinoMoto/ryu + Claude)

Downloads last month
17
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support