# Reproducing the quantization Prerequisites: Linux, Docker with NVIDIA GPU support, an explicitly reserved GPU, sufficient CPU RAM/disk for source weights and sequential caches, and network access to the pinned public Hugging Face datasets. Calibration stores the BF16 model on CPU and sequentially onloads decoder layers onto the reserved GPU; this is separate from inference. Resource sizing and hardware can affect runtime and reproducibility. CUDA algorithms may produce small numerical differences even with the recorded random seed. The published manifest identifies the exact released bytes. The original run completed calibration of all 28 decoder layers at 2026-10-09 01:11:45 UTC, before packing/export auditing. It used local immutable quantizer image `sha256:e4f0dbe60e59050796a15b7b0c4646130a93acfd4e2431cde05923e401470b10`. This is an image identity receipt, not a public registry download reference. Record the separately built image ID and frozen source hashes on reproduction. The quantizer is CUDA-dependent: use the actual CUDA AWQ preflight, not only the data-free CPU packing check. The original calibration required bounded memory-limit adjustments, finishing with a 52 GiB physical RAM limit and a 72 GiB combined RAM/swap limit. Plan enough host reserve for other workloads; the illustrative Docker commands below do not provision swap or change host resource limits. The run's sampled GPU maximum must be reported with its sampling interval and background usage, not presented as an exact allocator maximum or a serving-capacity measurement. Run from this `reproduction/` directory. Paths and GPU device 0 below are illustrative runtime inputs; confirm the reserved GPU before launching Docker. Do not use the quantizer venv for serving. ```bash hf download JetBrains/Mellum2.1-12B-A2.5B-Thinking \ --revision 92ddae9fc7665e9f801d141d2e5a6b2caf2460c4 --local-dir source python - <<'PY' import json from pathlib import Path Path('source/source-provenance.json').write_text(json.dumps({ 'model_id': 'JetBrains/Mellum2.1-12B-A2.5B-Thinking', 'revision': '92ddae9fc7665e9f801d141d2e5a6b2caf2460c4'}) + '\n') PY docker build -f Dockerfile.quantize -t mellum21-quantizer:local . docker run --rm --network none --runtime runc -e NVIDIA_VISIBLE_DEVICES=void \ --entrypoint python mellum21-quantizer:local /opt/mellum/check_quantizer.py docker run --rm --gpus device=0 --entrypoint python \ mellum21-quantizer:local /opt/mellum/check_quantizer.py --awq docker run --rm --gpus device=0 --shm-size 20g \ -v "$PWD/source:/source:ro" -v "$PWD:/output" \ -v "$PWD/model-cache:/root/.cache/huggingface" \ mellum21-quantizer:local --source-dir /source --output-dir /output/model \ --num-samples 256 --max-seq-length 2048 --seed 1234 ``` The script refuses an existing output directory. Preserve failed exports under distinct names; do not overwrite a checkpoint used for comparison. Review `model/quantization-evidence/export-audit.json` and rerun the audit if needed: ```bash docker run --rm --network none --entrypoint python \ -v "$PWD/source:/source:ro" -v "$PWD/model:/model:ro" \ mellum21-quantizer:local /opt/mellum/audit_export.py /model --source-dir /source ``` A passing audit proves export structure and scale/head checks. Validate actual model loading, semantic output, schemas/tool calls, context lengths and matched runtime settings separately before using a reproduced checkpoint. For the tested serving installation, use the separately supplied `../serving/serve_compat.py` and [serving instructions](../serving/README.md). The launcher uses vLLM's supported pre-validation callable and unchanged tokenizer bytes to avoid its Mellum configuration compatibility failures. It was validated after calibration and is not part of the frozen quantizer sources. Do not apply it to calibration, patch upstream packages, rewrite the checkpoint configuration or substitute a CLI dictionary override.