{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# Qwen3-8B in 2 bits\n", "\n", "Weights stay **packed as bit-planes** for the whole run — never expanded to dense fp16.\n", "\n", "| | packed | dense fp16 |\n", "|---|---|---|\n", "| weights resident | **3.65 GiB** | 15.3 GiB |\n", "| 1 sequence, RTX 4090 | **322 tok/s** | 59 |\n", "| 16 sequences, RTX 4090 | **3452 tok/s** | 890 |\n", "| Apple M4, 1 sequence | **31 tok/s** | *does not fit* |\n", "\n", "The notebook works out which machine you are on and picks the kernels." ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 1. Weights" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from huggingface_hub import snapshot_download\n", "MODEL = snapshot_download('gitarist/Qwen3-8B-BPDQ-2bit')\n", "MODEL" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 2. Dependencies for this machine" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "import importlib.util, subprocess, sys\n", "import torch\n", "\n", "need = [p for p in ('transformers', 'accelerate') if importlib.util.find_spec(p) is None]\n", "if torch.cuda.is_available() and importlib.util.find_spec('vllm') is None:\n", " need.append('vllm==0.28.0') # CUDA only\n", "if need:\n", " subprocess.run([sys.executable, '-m', 'pip', 'install', '-q', *need], check=True)\n", "print('installed:', need or 'nothing needed')" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 3. Load" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "import sys; sys.path.insert(0, MODEL)\n", "import bpdq\n", "\n", "DEVICE = bpdq.pick_device() # cuda | mps | cpu\n", "ENGINE = 'vllm' if (DEVICE == 'cuda' and bpdq._VLLM) else 'torch'\n", "print(f'{DEVICE} / {ENGINE}')" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from transformers import AutoTokenizer\n", "tok = AutoTokenizer.from_pretrained(MODEL)\n", "tok.padding_side = 'left'\n", "if tok.pad_token is None:\n", " tok.pad_token = tok.eos_token\n", "\n", "if ENGINE == 'vllm':\n", " from vllm import LLM\n", " _llm = LLM(model=MODEL, quantization='bpdq', dtype='float16', max_model_len=4096,\n", " max_num_seqs=64, reasoning_parser='qwen3')\n", "\n", " def generate(prompts, max_tokens=512, think=False, think_budget=256):\n", " sp = bpdq.Sampling.for_mode(think, think_budget=think_budget if think else None)\n", " return [o.outputs[0].text for o in _llm.generate(prompts, sp.vllm(max_tokens))]\n", "else:\n", " _model, tok = bpdq.load(MODEL) # ~40 s: the repack runs once, on device\n", "\n", " def generate(prompts, max_tokens=512, think=False, think_budget=None):\n", " enc = tok(prompts, return_tensors='pt', padding=True).to(_model.device)\n", " out, _s = bpdq.decode(_model, enc.input_ids, max_tokens,\n", " attention_mask=enc.attention_mask,\n", " sampling=bpdq.Sampling.for_mode(think))\n", " return [tok.decode(o, skip_special_tokens=True) for o in out]\n", "\n", "print('ready')" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 4. Chat\n", "\n", "Qwen3's published sampling settings per mode. `think=True` reasons first; on vLLM\n", "`think_budget` closes the thought, since this model rarely stops on its own." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "history = []\n", "\n", "def chat(message, think=False, max_tokens=512):\n", " history.append({'role': 'user', 'content': message})\n", " prompt = tok.apply_chat_template(history, tokenize=False,\n", " add_generation_prompt=True, enable_thinking=think)\n", " reply = generate([prompt], max_tokens, think=think)[0]\n", " history.append({'role': 'assistant', 'content': reply})\n", " return reply\n", "\n", "def reset():\n", " history.clear()\n", "\n", "print(chat('In two sentences, what is a bit-plane?'))" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "print(chat('Now give an example.')) # the history carries over" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 5. Many at once" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "for r in generate(['Name three uses for a hash table.',\n", " 'What is 17 * 23?',\n", " 'Write a haiku about bit-planes.'], max_tokens=128):\n", " print(r.strip(), '\\n---')" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "import time\n", "for B in (1, 4, 16):\n", " prompts = ['Count slowly from one to fifty.'] * B\n", " generate(prompts, 8)\n", " t0 = time.time(); generate(prompts, 64); dt = time.time() - t0\n", " print(f'{B:3d} sequences: {B * 64 / dt:8.1f} tok/s')" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 6. Interactive" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "reset()\n", "while True:\n", " msg = input('you: ')\n", " if not msg.strip():\n", " break\n", " print('bot:', chat(msg), '\\n')" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 7. Check it\n", "\n", "Kernels against a dense reconstruction, then the real weights layer by layer, then\n", "batched decode against `model.generate`." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "!python {MODEL}/bpdq.py selftest {MODEL}" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "!python {MODEL}/bpdq.py bench {MODEL}" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "name": "python", "version": "3.11" } }, "nbformat": 4, "nbformat_minor": 5 }