# JevBench reproduction This package targets JevBench `v1.3.0` / commit `51a8d73fa798aa337bb1b26abd10995c0ab847e9` and the frozen Nimble encoder at commit `f136b3f75721fda4ea961f73993cc50b08488835`. The adapter performs one forward pass per decision and returns a softmax over the exact declared candidates. It generates zero tokens, does not truncate, and refuses inputs beyond the configured 8,192-token semantic limit. The dummy `reference.target` required by Nimble's inference record encoder is always the first declared label and is never inserted into the prompt or used to compute logits. ## Environment ```bash git clone https://github.com/fstandhartinger/jevbench.git cd jevbench git checkout 51a8d73fa798aa337bb1b26abd10995c0ab847e9 git clone https://github.com/bespokelabsai/nimble.git ../nimble git -C ../nimble checkout f136b3f75721fda4ea961f73993cc50b08488835 python -m venv .venv . .venv/bin/activate pip install torch transformers==5.17.0 peft==0.21.0 accelerate==1.15.0 \ huggingface-hub==1.32.0 safetensors export PYTHONPATH="$PWD/../nimble:$PYTHONPATH" huggingface-cli download jsaurabh/qwen3.5-9b-jev-data-mix-v2 \ bench/qwen35_nimble_lora.py bench/jevbench-registration.patch \ --local-dir /tmp/qwen35-nimble-submission cp /tmp/qwen35-nimble-submission/bench/qwen35_nimble_lora.py jevbench/adapters/ git apply /tmp/qwen35-nimble-submission/bench/jevbench-registration.patch ``` Install `causal-conv1d` and `flash-linear-attention` when supported for optimized Qwen3.5 kernels. The generic PyTorch fallbacks are correct but slower. ## Run ```bash JEVBENCH_WARM_LOAD=1 python -m jevbench.cli run \ --tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \ --adapter qwen35_nimble_lora \ --endpoint jsaurabh/qwen3.5-9b-jev-data-mix-v2 \ --revision MODEL_REPOSITORY_REVISION \ --model qwen3.5-9b-jev-data-mix-v2 \ --cost-basis self_hosted_gpu \ --results /tmp/qwen35-nimble/results.jsonl \ --manifest /tmp/qwen35-nimble/manifest.json \ --ledger /tmp/qwen35-nimble/ledger.jsonl \ --raw-dir /tmp/qwen35-nimble/raw ``` Replace `MODEL_REPOSITORY_REVISION` with the pinned Hugging Face commit listed in the benchmark request. ## Public reference result The frozen 231 public tasks produced 184/231 correct (79.65%), with 48/48 easy, 68/72 standard, and 68/111 hard. All 231 distributions were strict-valid with no failures. Brier was 0.2895 and ECE was 0.0892. The complete public records are in `nimble-data-mix-v2-jevbench-public.json` at the repository root. The public run used an A100 40 GB in bf16 with the generic correct PyTorch implementations for Qwen3.5's causal convolution and gated delta rule. The evaluator should measure latency on its own selected hardware and apply JevBench's normal self-hosted adjustment and cost policy.