# Serving compatibility launcher `serve_compat.py` is project-owned serving glue, independent of the frozen calibration sources. Its SHA-256 and the tested serving image/version appear in `../publication-manifest.json`. The tested Ampere fork uses vLLM 0.29.0 and Transformers 5.17.0, with local immutable image `sha256:68d4a7a5303358be89f9cc6e035fa4d719a5fac554d912e6b31d16998b35a268`. This is an identity receipt, not a public image download. Compatibility with other installations has not been established. That installation's MellumConfig inherits Qwen3MoeConfig, whose constructor adds an unused outer scalar to the source's nested RoPE parameters. The supported callable `hf_overrides` restores the exact checkpoint dictionary before validation. Full-attention layers retain the source YaRN factor 16, theta 500000, original context 8192, beta 32/1 and attention factor 1.2772588722239782; sliding-attention layers retain default RoPE with theta 500000. The tokenizer registry reads configuration independently, so the launcher supplies a temporary tokenizer-only directory with original file bytes and `--tokenizer-mode hf`. That directory remains alive until serving ends. No checkpoint files or installed packages are changed. A CLI dictionary override is applied too late and must not replace this callable launcher. 1. Reserve a GPU and enter the compatible serving environment, which must provide vLLM, Transformers, uvloop and the tested attention/quantization kernels. Download this repository at its immutable published revision as shown in the model card. This includes `serving/serve_compat.py`. 2. Run a config/tokenizer preflight from the downloaded repository root: ```bash VLLM_MARLIN_INPUT_DTYPE=int8 python serving/serve_compat.py \ --model . --compat-config-only --quantization compressed-tensors \ --dtype bfloat16 --max-model-len 131072 --kv-cache-dtype auto ``` Expect `status: passed`, 28 decoder layers, the exact checkpoint RoPE dictionaries and source BOS/EOS IDs 0/28. This uses actual engine config and tokenizer registry construction; it does not load weights or prove model quality. Missing core tokenizer files, the wrong model/layer layout, or changed RoPE semantics must fail. The launcher owns tokenizer selection and overrides: omit `--tokenizer` and `--hf-overrides`. 3. Use the model-card launch command for actual serving. Confirm model loading, selected attention/MoE kernels, tokenizer IDs, thinking control and the allocated KV capacity. Run semantic/schema/tool tests and the context lengths needed by the workload before drawing quality or performance conclusions. FP8 KV needs its own matched comparison; the artifact does not calibrate KV scales. Preserve failures rather than automatically restarting the process. Actual preflight passed for both the pinned BF16 source and quantized configuration in the tested installation. The BF16 source also fully loaded with TP2 and completed a short generation. The derivative's completed quality/performance evidence belongs in the model card; these source-control and compatibility checks do not establish that evidence.