jsaurabh's picture
Add pinned JevBench adapter and submission package
af0d510 verified
|
Raw History Blame Contribute Delete
2.83 kB

JevBench reproduction

This package targets JevBench v1.3.0 / commit 51a8d73fa798aa337bb1b26abd10995c0ab847e9 and the frozen Nimble encoder at commit f136b3f75721fda4ea961f73993cc50b08488835.

The adapter performs one forward pass per decision and returns a softmax over the exact declared candidates. It generates zero tokens, does not truncate, and refuses inputs beyond the configured 8,192-token semantic limit. The dummy reference.target required by Nimble's inference record encoder is always the first declared label and is never inserted into the prompt or used to compute logits.

Environment

git clone https://github.com/fstandhartinger/jevbench.git
cd jevbench
git checkout 51a8d73fa798aa337bb1b26abd10995c0ab847e9

git clone https://github.com/bespokelabsai/nimble.git ../nimble
git -C ../nimble checkout f136b3f75721fda4ea961f73993cc50b08488835

python -m venv .venv
. .venv/bin/activate
pip install torch transformers==5.17.0 peft==0.21.0 accelerate==1.15.0 \
  huggingface-hub==1.32.0 safetensors
export PYTHONPATH="$PWD/../nimble:$PYTHONPATH"

huggingface-cli download jsaurabh/qwen3.5-9b-jev-data-mix-v2 \
  bench/qwen35_nimble_lora.py bench/jevbench-registration.patch \
  --local-dir /tmp/qwen35-nimble-submission
cp /tmp/qwen35-nimble-submission/bench/qwen35_nimble_lora.py jevbench/adapters/
git apply /tmp/qwen35-nimble-submission/bench/jevbench-registration.patch

Install causal-conv1d and flash-linear-attention when supported for optimized Qwen3.5 kernels. The generic PyTorch fallbacks are correct but slower.

Run

JEVBENCH_WARM_LOAD=1 python -m jevbench.cli run \
  --tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \
  --adapter qwen35_nimble_lora \
  --endpoint jsaurabh/qwen3.5-9b-jev-data-mix-v2 \
  --revision MODEL_REPOSITORY_REVISION \
  --model qwen3.5-9b-jev-data-mix-v2 \
  --cost-basis self_hosted_gpu \
  --results /tmp/qwen35-nimble/results.jsonl \
  --manifest /tmp/qwen35-nimble/manifest.json \
  --ledger /tmp/qwen35-nimble/ledger.jsonl \
  --raw-dir /tmp/qwen35-nimble/raw

Replace MODEL_REPOSITORY_REVISION with the pinned Hugging Face commit listed in the benchmark request.

Public reference result

The frozen 231 public tasks produced 184/231 correct (79.65%), with 48/48 easy, 68/72 standard, and 68/111 hard. All 231 distributions were strict-valid with no failures. Brier was 0.2895 and ECE was 0.0892. The complete public records are in nimble-data-mix-v2-jevbench-public.json at the repository root.

The public run used an A100 40 GB in bf16 with the generic correct PyTorch implementations for Qwen3.5's causal convolution and gated delta rule. The evaluator should measure latency on its own selected hardware and apply JevBench's normal self-hosted adjustment and cost policy.