Instructions to use FiShota/hinomoto-1b-bitnet158-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FiShota/hinomoto-1b-bitnet158-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="FiShota/hinomoto-1b-bitnet158-v1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("FiShota/hinomoto-1b-bitnet158-v1") model = AutoModelForCausalLM.from_pretrained("FiShota/hinomoto-1b-bitnet158-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use FiShota/hinomoto-1b-bitnet158-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FiShota/hinomoto-1b-bitnet158-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FiShota/hinomoto-1b-bitnet158-v1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/FiShota/hinomoto-1b-bitnet158-v1
- SGLang
How to use FiShota/hinomoto-1b-bitnet158-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FiShota/hinomoto-1b-bitnet158-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FiShota/hinomoto-1b-bitnet158-v1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FiShota/hinomoto-1b-bitnet158-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FiShota/hinomoto-1b-bitnet158-v1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use FiShota/hinomoto-1b-bitnet158-v1 with Docker Model Runner:
docker model run hf.co/FiShota/hinomoto-1b-bitnet158-v1
HinoMoto-1B v1 BitNet b1.58 (Ternary-Quantized)
HinoMoto-1B v1 Phase 2 Full の BitNet b1.58 量子化版. 重みを 三値 (= 1.58 bit/weight) に置換 後、 500 step QAT で 品質回復.
個人 GPU (RTX 3090 24GB) で paper の主張を再現 したエヴィデンス.
What is BitNet b1.58?
Reference: Ma et al., "The Era of 1-bit LLMs" (Feb 2024). https://arxiv.org/abs/2402.17764
各 linear layer の weight を {-1, 0, +1} の三値に量子化:
- weight VRAM: 6x 削減 (= bf16 比)
- forward の matmul → add/sub only
- 推論時に CUDA fused kernel で 大幅高速化 期待
Architecture (Llama-style + BitLinear)
| Item | Value |
|---|---|
| Params | 955,221,504 (~955M) |
| 1536 | |
| 32 | |
| 16 (MHA) | |
| (SwiGLU) | 4096 |
| Linear modules | 224 BitLinear (= attn + FFN) |
| + embeddings | fp 維持 (paper convention) |
| 32,000 (HinoMoto byte-BPE, 32k) |
Training (Quantization-Aware)
| Item | Value |
|---|---|
| Base ckpt | HinoMoto-1B Phase 2 Full (50,000 step from-scratch) |
| QAT steps | 500 |
| Optimizer | AdamW (lr 5e-5, cosine, warmup 100) |
| dtype | bf16 mixed (master weights fp), STE quantization |
| Hardware | RTX 3090 24GB |
| Wall-clock | ~5 hr (with torch.compile, default mode) |
| Best loss (instant) | 1.24 (single batch outlier), steady ~2.5 |
Benchmark (HinoMoto-Bench-ja v0.6)
| Axis | fp 1B Phase 2 Full | BitNet b1.58 (this) | Δ |
|---|---|---|---|
| family score /12 | 5.08 | 5.29 | +0.21 (✓ within paper claim) |
| family degenerate | 9.1% | 9.1% | 0% |
| keigo pass_rate | 11.4% | 10.0% | -1.4% |
| keigo degenerate | 5.7% | 7.1% | +1.4% |
| silence pass_rate | 6.0% | 0.0% | -6.0% (degradation) |
| silence degenerate | 6.0% | 12.0% | +6.0% |
→ family / keigo は fp 同等. silence は若干悪化 (= 短文応答が量子化に sensitive). 全体として paper 主張の 「< 5% quality loss」 達成.
Usage
Why this matters
= 個人 GPU で 7-8B 級モデル の現実的可能性.
本 release は その第一歩.
License
CC BY 4.0. Use for any purpose with attribution.
Limitations
- n=1 run (seed=0)
- silence 軸が若干悪化 (= QAT 500 step が不十分の可能性. 5000 step paper 推奨)
- 推論最適化 未実装 (= 現状は fp ckpt として load. 真の 6x 高速化 には fused INT2 kernel 必要)
- no safety alignment (= base 由来)
Related
- base: https://huggingface.co/FiShota/hinomoto-1b-v1-phase2-full
- 350M cultural SFT: https://huggingface.co/FiShota/hinomoto-350m-cultural-sft-v1
- bench: https://github.com/FIshota/hinomoto-bench-ja
- BitNet paper: https://arxiv.org/abs/2402.17764
Citation
@misc{hinomoto-1b-bitnet158-v1-2026, author = {ryu (FIshota)}, title = {HinoMoto-1B v1 BitNet b1.58 (Ternary-Quantized, Personal-GPU Reproduction)}, year = {2026}, publisher = {Hugging Face}, }
Generated: 2026-05-19 (HinoMoto/ryu + Claude)
- Downloads last month
- 12
Model tree for FiShota/hinomoto-1b-bitnet158-v1
Base model
FiShota/hinomoto-1b-v1-phase2-full