- HinoMoto-100M v15 (WSD + z-loss + EMA, seed=0)
HinoMoto-100M v15 (WSD + z-loss + EMA, seed=0)
100M parameter Japanese language model trained from scratch with WSD + z-loss schedule and EMA decay=0.999.
About the developer โ One person (FiShota) building a Japanese LM stack from scratch on a single RTX 3090. HinoMoto = from-scratch JP LM family, Yamato = legal/admin SFT specialist, HinoMoto-Bench-ja = cultural-axis benchmark. Honest research: every release links to a write-up that includes the negative results. GitHub ยท Bench ยท X
Quick facts
- Parameters: 100M (12 layer, d_model=512, vocab=9506)
- Tokenizer: Byte-BPE 9506 tokens (balanced corpus)
- Training: 20000 steps, RTX 3090 24GB, fp32, ~28 min
- Schedule: WSD (Warmup-Stable-Decay), warmup=400, stable=80%, decay=20%
- Regularizer: z-loss coef=1e-4
- EMA: decay=0.999 (CPU shadow)
PPL (51 / 200 / 789 prompts)
| Eval set | PPL |
|---|---|
| 51 short prompts | 15.94 |
| 200 mixed prompts | 32.08 |
| 789 broad prompts | 48.51 |
Honest comparison with v12 (WSD seed=2) and cosine baselines
3-seed ร 2-schedule ร 3 eval-set ่จๆธฌ:
| Schedule | seed=0 | seed=1 | seed=2 | mean (51) | std (51) | mean (789) | std (789) |
|---|---|---|---|---|---|---|---|
| Cosine | v7 (17.72) | v8 (17.09) | v14 (16.58) | 17.13 | 0.57 | 48.03 | 0.18 |
| WSD+z-loss | v15 (15.94) | v11/v16 (16.50) | v12 (16.35) | 16.26 | 0.29 | 48.76 | 0.45 |
paired t-test (cosine vs WSD): p=0.148 (51), p=0.484 (200), p=0.534 (789) = ๅ จ eval ใงๆๆๅทฎใชใ.
โ WSD โ cosine at 100M JP PPL (no statistically significant difference).
Update โ 2026-05-09 โ Phase 1 (350M scale-up) progress
The HinoMoto pipeline has scaled up to 350M params with the same architecture family (d_model=1024, 24 layers, vocab=9506).
| Run | Params | Schedule | Steps | Final ppl | Notes |
|---|---|---|---|---|---|
| 100M v15 (this card) | 100M (43M excl embed) | WSD+z-loss+EMA | 20,000 | 15.94 (51-prompt) | this card |
| 350M smoke | 318M | WSD+z-loss+EMA bf16 | 5,000 | 9.53 (training) | architecture verified |
| 350M Phase 1 full | 318M | WSD+z-loss+EMA bf16 | 50,000 | (in progress) | RTX 3090 single GPU, ~22h |
Architecture scales cleanly from 100M โ 350M with the same recipe. Next: HinoMoto-1B (Phase 2) after corpus expansion (78MB โ 5GB+ planned).
โ Detailed retro: see HinoMoto ้็บใใผใ #10 (post Phase 1 ๅฎไบๅพ).
How to use
import torch
from hinomoto.model.hinomoto_model import HinoMotoModel, HinoMotoConfig
from hinomoto.tokenizer.byte_bpe import ByteBPETokenizer
ck = torch.load("ckpt_ema_step_020000_final.pt", map_location="cpu", weights_only=False)
cfg = HinoMotoConfig(**ck["config"])
model = HinoMotoModel(cfg)
model.load_state_dict(ck["shadow"], strict=False)
model.lm_head.weight = model.tok_embed.weight # tied
model.eval()
tok = ByteBPETokenizer.load("tokenizer.json")
ids = tok.encode("ๆฅๆฌ่ชใฎๆ็ซ ใขใใซใไฝใใซใฏ")
# ... your inference loop
Training command
python -m hinomoto.train.train_lm \
--tokenizer tokenizer_v3_32k_clean.json \
--corpus all_v6_balanced.txt \
--output-dir artifacts/smoke_100m_v15_wsd_zloss_ema \
--config configs/main_run_100m_v3.json \
--max-steps 20000 --warmup-steps 400 \
--batch-size 2 --grad-accum 4 --seq-len 512 \
--lr 3e-4 --device cuda \
--dtype fp32 --seed 0 \
--ema-decay 0.999 \
--lr-schedule wsd --wsd-decay-frac 0.2 \
--z-loss-coef 1e-4 \
--spike-detect
Reproducibility
The HinoMoto pipeline is fp32 bit-deterministic in our environment
(Windows 11 + WSL2 Ubuntu 24.04 + RTX 3090).
Same --seed โ same weights, even across days/runs (verified: v11 = v16 with seed=1).
License
Apache-2.0
Citation
Coming soon (arXiv WIP, see paper_drafts/hinomoto_arxiv_outline.md).
All FiShota models (full family map)
From-scratch JP LM (HinoMoto, Apache-2.0)
| Repo | Scale | Stage | Status |
|---|---|---|---|
| hinomoto-100m-v15-wsd-zloss-ema | 100M | research baseline (this card) | โ stable |
| hinomoto-100m-v12-wsd-zloss-seed2 | 100M | sister seed=2 (reproducibility) | โ stable |
Sarashina2.2-3B-instruct + cultural-axis SFT (MIT)
| Repo | Recipe | Notes |
|---|---|---|
| sarashina2.2-3b-sft-v3-Q4_K_M-gguf | 4-axis SFT, 4 quants (Q3/Q4/Q5/Q6) | family / keigo / silence / atmosphere |
| sarashina2.2-3b-sft-v4-gguf | sft_v3 + 11% NLI replay | catastrophic-forgetting repair |
| sarashina2.2-3b-sft-v4-dpo-gguf | sft_v4 + DPO 108 pairs | preference tuning |
Yamato โ Japanese legal / administrative specialist (MIT, LoRA r-ablation)
| Repo | LoRA r | Bench % | Notes |
|---|---|---|---|
| yamato-3b-v1-legal-gguf | 16 | 46.7% | baseline |
| yamato-3b-v2-legal-gguf | 16 | โ | 2nd pass |
| yamato-3b-v3-r32-legal-gguf | 32 | 48.9% | +2.2 |
| yamato-3b-v4-r64-legal-gguf | 64 | 54.3% | +5.4 |
| yamato-3b-v5-r128-legal-gguf โญ | 128 | 57.6% | best โ see dev-note #9 |
Companion repos (GitHub)
- hinomoto-bench-ja โ cultural-axis benchmark, CC BY 4.0, 418+ items
- hinomoto-model โ training pipeline (this card's recipe)
- hinomoto-mojo โ Mojo SIMD/parallel kernels for inference
- rtk โ Rust corpus-cleaning tool
๐ 350M Phase 1 training curve (live)
โ HinoMoto-350M Phase 1 (in-progress). Updated periodically. Left: perplexity (log scale). Right: loss (linear). See HinoMoto ้็บใใผใ #10 (post-completion retro).
