GGUF
Japanese
How to use from
Lemonade
Pull the model
# Download Lemonade from https://lemonade-server.ai/
lemonade pull WariHima/hourai3-yomi-90m-v1-gguf
Run and chat with the model
lemonade run user.hourai3-yomi-90m-v1-gguf-{{QUANT_TAG}}
List all available models
lemonade list
Quick Links

hourai3 yomi 90m v1

zenz v3互換 読み推定モデル

architecture:

  • lfm2moe

tarined by:

  • rtx3060 * 8m training code in ./training_codes

tokenizer from ku-nlp/gpt2-small-japanese-char
+(unused qwen3 special toknes from base model)

dataset:

  • Miwa-Keita/zenz-v2.5-dataset use wikipedia part 40k rows 1 epoch

qunat type

  • BF16

example code (python)

import llama_cpp

llm =  llama_cpp.Llama(
     model_path="./hourai3-yomi-90m-v1.gguf",",
        embedding=False,
        verbose=False
    )

while True:
      input_text = input("入力した文章の読みが出力されます。exitを入力で終了:")
    
    if input_text == "exit":
        break
        
    text = f"<s>\uEE00{input_text}\uEE01"
    
    output_text = llm(
        text,
        max_tokens=10,          # 生成する最大トークン数
        temperature=0.0,       # ランダム性(0.0に近づくほど確実な出力、1.0以上で多様化)
        top_k=40,               # 上位k個の候補に絞り込む
        top_p=0.95,             # 累積確率p以下の候補に絞り込む
        repeat_penalty=1.1,     # 同じ単語の繰り返しを抑制するペナルティ
        stop=[],
        echo=False
    )["choices"][0]["text"]
    print(f"入力|出力: {input_text}|{output_text}")

convert

変換時、qwen3nextがgpt-2トークナイザで使用することを想定されていなかったため、
llama.cppリポジトリのconversion/base.pyファイルの
以下のエラーをバイパスする必要がありました。
1687=1699付近
        if res is None:
            logger.warning("\n")
            logger.warning("**************************************************************************************")
            logger.warning("** WARNING: The BPE pre-tokenizer was not recognized!")
            logger.warning("**          There are 2 possible reasons for this:")
            logger.warning("**          - the model has not been added to convert_hf_to_gguf_update.py yet")
            logger.warning("**          - the pre-tokenization config has changed upstream")
            logger.warning("**          Check your model files and convert_hf_to_gguf_update.py and update them accordingly.")
            logger.warning("** ref:     https://github.com/ggml-org/llama.cpp/pull/6920")
            logger.warning("**")
            logger.warning(f"** chkhsh:  {chkhsh}")
            logger.warning("**************************************************************************************")
            logger.warning("\n")
            
            return "default"
            #raise NotImplementedError("BPE pre-tokenizer was not recognized - update get_vocab_base_pre()")
想定されていないことによる変換時のエラーなので、推論時はmainstreamのllama.cppで動作します。
Downloads last month
37
GGUF
Model size
84.9M params
Architecture
lfm2moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WariHima/hourai3-yomi-90m-v1-gguf

Quantized
(1)
this model

Collection including WariHima/hourai3-yomi-90m-v1-gguf