HinoMoto-100M v15 (WSD + z-loss + EMA, seed=0)

100M parameter Japanese language model trained from scratch with WSD + z-loss schedule and EMA decay=0.999.

About the developer โ€” One person (FiShota) building a Japanese LM stack from scratch on a single RTX 3090. HinoMoto = from-scratch JP LM family, Yamato = legal/admin SFT specialist, HinoMoto-Bench-ja = cultural-axis benchmark. Honest research: every release links to a write-up that includes the negative results. GitHub ยท Bench ยท X

Quick facts

  • Parameters: 100M (12 layer, d_model=512, vocab=9506)
  • Tokenizer: Byte-BPE 9506 tokens (balanced corpus)
  • Training: 20000 steps, RTX 3090 24GB, fp32, ~28 min
  • Schedule: WSD (Warmup-Stable-Decay), warmup=400, stable=80%, decay=20%
  • Regularizer: z-loss coef=1e-4
  • EMA: decay=0.999 (CPU shadow)

PPL (51 / 200 / 789 prompts)

Eval set PPL
51 short prompts 15.94
200 mixed prompts 32.08
789 broad prompts 48.51

Honest comparison with v12 (WSD seed=2) and cosine baselines

3-seed ร— 2-schedule ร— 3 eval-set ่จˆๆธฌ:

Schedule seed=0 seed=1 seed=2 mean (51) std (51) mean (789) std (789)
Cosine v7 (17.72) v8 (17.09) v14 (16.58) 17.13 0.57 48.03 0.18
WSD+z-loss v15 (15.94) v11/v16 (16.50) v12 (16.35) 16.26 0.29 48.76 0.45

paired t-test (cosine vs WSD): p=0.148 (51), p=0.484 (200), p=0.534 (789) = ๅ…จ eval ใงๆœ‰ๆ„ๅทฎใชใ—.

โ†’ WSD โ‰ˆ cosine at 100M JP PPL (no statistically significant difference).

Update โ€” 2026-05-09 โ€” Phase 1 (350M scale-up) progress

The HinoMoto pipeline has scaled up to 350M params with the same architecture family (d_model=1024, 24 layers, vocab=9506).

Run Params Schedule Steps Final ppl Notes
100M v15 (this card) 100M (43M excl embed) WSD+z-loss+EMA 20,000 15.94 (51-prompt) this card
350M smoke 318M WSD+z-loss+EMA bf16 5,000 9.53 (training) architecture verified
350M Phase 1 full 318M WSD+z-loss+EMA bf16 50,000 (in progress) RTX 3090 single GPU, ~22h

Architecture scales cleanly from 100M โ†’ 350M with the same recipe. Next: HinoMoto-1B (Phase 2) after corpus expansion (78MB โ†’ 5GB+ planned).

โ†’ Detailed retro: see HinoMoto ้–‹็™บใƒŽใƒผใƒˆ #10 (post Phase 1 ๅฎŒไบ†ๅพŒ).

How to use

import torch
from hinomoto.model.hinomoto_model import HinoMotoModel, HinoMotoConfig
from hinomoto.tokenizer.byte_bpe import ByteBPETokenizer

ck = torch.load("ckpt_ema_step_020000_final.pt", map_location="cpu", weights_only=False)
cfg = HinoMotoConfig(**ck["config"])
model = HinoMotoModel(cfg)
model.load_state_dict(ck["shadow"], strict=False)
model.lm_head.weight = model.tok_embed.weight  # tied
model.eval()

tok = ByteBPETokenizer.load("tokenizer.json")
ids = tok.encode("ๆ—ฅๆœฌ่ชžใฎๆ–‡็ซ ใƒขใƒ‡ใƒซใ‚’ไฝœใ‚‹ใซใฏ")
# ... your inference loop

Training command

python -m hinomoto.train.train_lm \
  --tokenizer tokenizer_v3_32k_clean.json \
  --corpus all_v6_balanced.txt \
  --output-dir artifacts/smoke_100m_v15_wsd_zloss_ema \
  --config configs/main_run_100m_v3.json \
  --max-steps 20000 --warmup-steps 400 \
  --batch-size 2 --grad-accum 4 --seq-len 512 \
  --lr 3e-4 --device cuda \
  --dtype fp32 --seed 0 \
  --ema-decay 0.999 \
  --lr-schedule wsd --wsd-decay-frac 0.2 \
  --z-loss-coef 1e-4 \
  --spike-detect

Reproducibility

The HinoMoto pipeline is fp32 bit-deterministic in our environment (Windows 11 + WSL2 Ubuntu 24.04 + RTX 3090). Same --seed โ†’ same weights, even across days/runs (verified: v11 = v16 with seed=1).

License

Apache-2.0

Citation

Coming soon (arXiv WIP, see paper_drafts/hinomoto_arxiv_outline.md).

All FiShota models (full family map)

From-scratch JP LM (HinoMoto, Apache-2.0)

Repo Scale Stage Status
hinomoto-100m-v15-wsd-zloss-ema 100M research baseline (this card) โœ… stable
hinomoto-100m-v12-wsd-zloss-seed2 100M sister seed=2 (reproducibility) โœ… stable

Sarashina2.2-3B-instruct + cultural-axis SFT (MIT)

Repo Recipe Notes
sarashina2.2-3b-sft-v3-Q4_K_M-gguf 4-axis SFT, 4 quants (Q3/Q4/Q5/Q6) family / keigo / silence / atmosphere
sarashina2.2-3b-sft-v4-gguf sft_v3 + 11% NLI replay catastrophic-forgetting repair
sarashina2.2-3b-sft-v4-dpo-gguf sft_v4 + DPO 108 pairs preference tuning

Yamato โ€” Japanese legal / administrative specialist (MIT, LoRA r-ablation)

Repo LoRA r Bench % Notes
yamato-3b-v1-legal-gguf 16 46.7% baseline
yamato-3b-v2-legal-gguf 16 โ€” 2nd pass
yamato-3b-v3-r32-legal-gguf 32 48.9% +2.2
yamato-3b-v4-r64-legal-gguf 64 54.3% +5.4
yamato-3b-v5-r128-legal-gguf โญ 128 57.6% best โ€” see dev-note #9

Companion repos (GitHub)

  • hinomoto-bench-ja โ€” cultural-axis benchmark, CC BY 4.0, 418+ items
  • hinomoto-model โ€” training pipeline (this card's recipe)
  • hinomoto-mojo โ€” Mojo SIMD/parallel kernels for inference
  • rtk โ€” Rust corpus-cleaning tool

๐Ÿ“Š 350M Phase 1 training curve (live)

350M training curve

โ†‘ HinoMoto-350M Phase 1 (in-progress). Updated periodically. Left: perplexity (log scale). Right: loss (linear). See HinoMoto ้–‹็™บใƒŽใƒผใƒˆ #10 (post-completion retro).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support