GGUF
Japanese
How to use from
Unsloth Studio
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for WariHima/hourai3-zenz-90m-v1-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for WariHima/hourai3-zenz-90m-v1-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for WariHima/hourai3-zenz-90m-v1-gguf to start chatting
Quick Links

hourai2 zenz 90m v1 gguf

zenz v3互換のカナ漢字変換モデルです
can use upsteream llama.cpp

architecture:

  • lfm2moe

tarined by:

  • rtx3060 * 8m
    tokenizer from ku-nlp/gpt2-small-japanese-char
    (+ qwen3 special toknes from base model)

dataset:

  • Miwa-Keita/zenz-v2.5-dataset use wikipedia part 40k rows 1 epoch

qunat type

  • BF16

example code (python)

import llama_cpp

llm =  llama_cpp.Llama(
     model_path="./hourai3-zenz-90m-v1.gguf",
        embedding=False,
        verbose=False
    )

while True:
    input_text = input("入力した文章が漢字交じり文に変換されます。カタカナのみで入力。exitを入力で終了:")
    
    if input_text == "exit":
        break
        
    text = f"<s>\uEE00{input_text}\uEE01"
    
    output_text = llm(
        text,
        max_tokens=10,          # 生成する最大トークン数
        temperature=0.0,       # ランダム性(0.0に近づくほど確実な出力、1.0以上で多様化)
        top_k=40,               # 上位k個の候補に絞り込む
        top_p=0.95,             # 累積確率p以下の候補に絞り込む
        repeat_penalty=1.1,     # 同じ単語の繰り返しを抑制するペナルティ
        stop=[],
        echo=False
    )["choices"][0]["text"]
    print(f"入力|出力: {input_text}|{output_text}")

convert

変換時、qwen3nextがgpt-2トークナイザで使用することを想定されていなかったため、
llama.cppリポジトリのconversion/base.pyファイルの
以下のエラーをバイパスする必要がありました。
1687=1699付近
        if res is None:
            logger.warning("\n")
            logger.warning("**************************************************************************************")
            logger.warning("** WARNING: The BPE pre-tokenizer was not recognized!")
            logger.warning("**          There are 2 possible reasons for this:")
            logger.warning("**          - the model has not been added to convert_hf_to_gguf_update.py yet")
            logger.warning("**          - the pre-tokenization config has changed upstream")
            logger.warning("**          Check your model files and convert_hf_to_gguf_update.py and update them accordingly.")
            logger.warning("** ref:     https://github.com/ggml-org/llama.cpp/pull/6920")
            logger.warning("**")
            logger.warning(f"** chkhsh:  {chkhsh}")
            logger.warning("**************************************************************************************")
            logger.warning("\n")
            
            return "default"
            #raise NotImplementedError("BPE pre-tokenizer was not recognized - update get_vocab_base_pre()")
想定されていないことによる変換時のエラーなので、推論時はmainstreamのllama.cppで動作します。
Downloads last month
170
GGUF
Model size
84.9M params
Architecture
lfm2moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including WariHima/hourai3-zenz-90m-v1-gguf