Instructions to use FiShota/hinomoto-350m-cultural-sft-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FiShota/hinomoto-350m-cultural-sft-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="FiShota/hinomoto-350m-cultural-sft-v1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("FiShota/hinomoto-350m-cultural-sft-v1") model = AutoModelForCausalLM.from_pretrained("FiShota/hinomoto-350m-cultural-sft-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use FiShota/hinomoto-350m-cultural-sft-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FiShota/hinomoto-350m-cultural-sft-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FiShota/hinomoto-350m-cultural-sft-v1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/FiShota/hinomoto-350m-cultural-sft-v1
- SGLang
How to use FiShota/hinomoto-350m-cultural-sft-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FiShota/hinomoto-350m-cultural-sft-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FiShota/hinomoto-350m-cultural-sft-v1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FiShota/hinomoto-350m-cultural-sft-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FiShota/hinomoto-350m-cultural-sft-v1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use FiShota/hinomoto-350m-cultural-sft-v1 with Docker Model Runner:
docker model run hf.co/FiShota/hinomoto-350m-cultural-sft-v1
Use Docker images
docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "FiShota/hinomoto-350m-cultural-sft-v1" \
--host 0.0.0.0 \
--port 30000# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "FiShota/hinomoto-350m-cultural-sft-v1",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'HinoMoto-350M Cultural-SFT v1
HinoMoto-350M (from-scratch Japanese LM) を HinoMoto-Bench-ja v0.6 で Cultural SFT した merged 版.
「家族・敬語・沈黙・思いやり・地方文化・ニュアンス・世代差・職場」 の 8 文化軸での応答に特化.
What this is
- base: HinoMoto-350M v1 (318M params, from-scratch 日本語 LM, vocab 9506)
- fine-tune: LoRA r=32 / alpha=64 / 5 epoch
- data: HinoMoto-Bench-ja v0.6 train split (551 items / 8 cultural axes)
- adapter merged into base (single safetensors file, no peft loading needed)
SFT 効果 (40 items eval, 8軸×5)
| Metric | base alone | +Cultural SFT (this model) |
|---|---|---|
| gold bigram overlap | 0.039 | 0.181 (+364%) |
| mean output length | 104 chars | 102 chars (適切に保持) |
| artifact rate | 14/40 (35%) | 1/40 (2.5%) |
| mean_token_accuracy (train end) | — | 86.1% |
| train_loss (final) | — | 0.78 |
→ from-scratch 350M base に対して gold overlap が 3.6x 改善 + artifact ほぼ除去.
Training summary
| Item | Value |
|---|---|
| Method | LoRA (r=32, alpha=64, target=all linears) |
| Epochs | 5 |
| Batch | 2 × grad_accum 4 (effective 8) |
| Seq len | 512 |
| LR | 3e-4, cosine schedule, warmup 10% |
| Precision | bf16 mixed |
| Hardware | RTX 3090 24GB |
| Wall-clock | ~10 min |
| Data | 551 train items (80% of bench v0.6) |
| Eval split | 141 held-out items |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"FiShota/hinomoto-350m-cultural-sft-v1",
dtype=torch.bfloat16,
)
tok = AutoTokenizer.from_pretrained("FiShota/hinomoto-350m-cultural-sft-v1")
prompt_template = """{axis}軸: {axis_def}
場面: {context}
ユーザー発言: {user}
回答:"""
# Example: keigo axis
prompt = prompt_template.format(
axis="keigo",
axis_def="敬語 (尊敬・謙譲・丁寧) の使い分けが問われる場面",
context="上司に報告メールを書く場面",
user="昨日の会議について",
)
inputs = tok(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=80, do_sample=False, top_k=1)
print(tok.decode(out[0], skip_special_tokens=True))
License
CC BY 4.0. Use for any purpose with attribution.
Intended Uses & Limitations
Intended uses:
- Research on Japanese cultural-axis evaluation
- Demonstration of from-scratch JP LM + LoRA SFT on consumer GPU
- Educational reference for small-LM bench-trained models
Out-of-scope:
- Production assistant / chatbot (no safety alignment, limited corpus)
- Tasks requiring broad world knowledge (78MB pre-training corpus is small)
- Code generation, math, factual QA outside the cultural-axis domain
Bias, Risks, Limitations
- n=1 run (seed=0 only) — variance not measured
- Bench-trained: works best on cultural-axis prompts in bench v0.6 format. For general chat, may be suboptimal.
- No safety alignment. May produce inappropriate content if prompted out-of-distribution.
- Limited corpus (350M base trained on 78MB; SFT on 551 items): expect mistakes & repetitions on out-of-distribution prompts.
- Japanese cultural perspective embedded: bench items reflect a specific (urban Japanese) view; minority dialects / regional variations underrepresented.
Transparency: Bench / Corpus Leak Audit
The training corpus and HinoMoto-Bench-ja have been audited for n-gram overlap (8-char):
| Bench axis | mean n-gram leak | max | items >30% leak |
|---|---|---|---|
| family | 2.7% | 51.8% | 2 / 110 |
| keigo | 22.9% | 77.8% | 25 / 70 (※ note below) |
| silence | 1.4% | 27.3% | 0 / 50 |
Note on keigo: high overlap is structural — top "leaks" are standard formulaic Japanese expressions like 「今しばらくお待ちください」「おはようございます」 that appear naturally in any Japanese corpus and are exactly the target output of the keigo evaluation axis. This is not contamination; it reflects the language's reliance on standardized polite forms.
The bench/corpus pair is considered publication-ready under this analysis.
Audit script: scripts/bench_leak_audit.py in the training repo.
Related
- bench data: https://github.com/FIshota/hinomoto-bench-ja
- base: https://huggingface.co/FiShota/hinomoto-350m-v1-llama (if published; otherwise see release notes)
- training repo: https://github.com/FIshota/hinomoto-model
Citation
@misc{hinomoto-350m-cultural-sft-2026,
author = {ryu (FIshota)},
title = {HinoMoto-350M Cultural-SFT v1},
year = {2026},
publisher = {Hugging Face},
}
Generated: 2026-05-13 (HinoMoto/ryu + Claude)
- Downloads last month
- 17
Install from pip and serve model
# Install SGLang from pip: pip install sglang# Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FiShota/hinomoto-350m-cultural-sft-v1" \ --host 0.0.0.0 \ --port 30000# Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FiShota/hinomoto-350m-cultural-sft-v1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'