SmolLM2-135M β€” 1BP format

HuggingFaceTB/SmolLM2-135M (apache-2.0) converted to 1BP β€” the single-file model format used by 1bit.systems's inference engine.

1BP packs everything a loader needs into one memory-mappable file: a 256-byte header with model config, a variable-length tensor index, and Q4NX-tiled (32Γ—256) 4-bit quantized weight data. No Python dependencies, no config files, no tokenizer files β€” just one file and it works.

Model Details

Property Value
Parameters 135M
Architecture llama
Quantization Q4NX (4-bit, 32Γ—256 tiles)
Source HuggingFaceTB/SmolLM2-135M
Format Single-file .1bp

Usage

# Download and run with npu_engine_universal
wget https://huggingface.co/bong-water-water-bong/SmolLM2-135M-1BP/resolve/main/SmolLM2-135M.1bp
./npu_engine_universal SmolLM2-135M.1bp 5
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for bong-water-water-bong/SmolLM2-135M-1BP

Finetuned
(929)
this model