How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf tdh111/bitnet-b1.58-2B-4T-GGUF
# Run inference directly in the terminal:
llama cli -hf tdh111/bitnet-b1.58-2B-4T-GGUF
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf tdh111/bitnet-b1.58-2B-4T-GGUF
# Run inference directly in the terminal:
llama cli -hf tdh111/bitnet-b1.58-2B-4T-GGUF
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf tdh111/bitnet-b1.58-2B-4T-GGUF
# Run inference directly in the terminal:
./llama-cli -hf tdh111/bitnet-b1.58-2B-4T-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf tdh111/bitnet-b1.58-2B-4T-GGUF
# Run inference directly in the terminal:
./build/bin/llama-cli -hf tdh111/bitnet-b1.58-2B-4T-GGUF
Use Docker
docker model run hf.co/tdh111/bitnet-b1.58-2B-4T-GGUF
Quick Links

The IQ2_BN and IQ2_BN_R4 version of microsoft/bitnet-b1.58-2B-4T-gguf for use with ik_llama.cpp.

I recommend the IQ2_BN_R4 version but you use -rtr on IQ2_BN to convert on runtime.

The chat template in the model looks incorrect (I did not change it, this is from the original Microsoft GGUF).

An example of correct usage from their transformers PR:

<|begin_of_text|>User: Hey, are you conscious? Can you talk to me?<|eot_id|>Assistant:

I was able to follow the example above and it worked for multi-turn conversations.

With the more general template (sourced from the paper) being:

<|begin_of_text|>System: {system_message}<|eot_id|> User: {user_message_1}<|eot_id|> Assistant: {assistant_message_1}<|eot_id|> User: {user_message_2}<|eot_id|> Assistant: {assistant_message_2}<|eot_id|>

Downloads last month
423
GGUF
Model size
3B params
Architecture
bitnet-25
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tdh111/bitnet-b1.58-2B-4T-GGUF

Quantized
(8)
this model