The reported accuracy could not be replicated.

#5
by DiyanGoodLuck - opened

I tried to replicate the results for GSM8K, IFEval, Humaneval+, and MBPP using default lm-eval config, but the numbers I got are way off from what's in the white paper. Is there a specific setup I can use to get results that match the white paper?

This is the command I use to start llama-cpp server:

    nohup ${LLAMA_CPP_DIR}/build/bin/llama-server \
    -m ${MODEL_PATH} \
    --port ${SERVER_PORT} \
    -c 4096 \
    --log-disable \
    > ${SERVER_LOG} 2>&1 &

IFEval, Humaneval+, and MBPP tasks are similar, but I did not set num_fewshot for them.

model: gguf
model_args:
  base_url: http://localhost:8080
tasks:
  - gsm8k
num_fewshot: 5
batch_size: 1
Tasks Version Filter n-shot Metric Value Stderr
humaneval_plus 1 create_test 0 pass@1 ↑ 0.2195 ± 0.0324
mbpp_plus 1 none 0 pass_at_1 ↑ 0.0000 ± 0.0000
gsm8k 3 flexible-extract 5 exact_match ↑ 0.8491 ± 0.0099
strict-match 5 exact_match ↑ 0.8506 ± 0.0098
ifeval 4 none 0 inst_level_loose_acc ↑ 0.4472 ± N/A
none 0 inst_level_strict_acc ↑ 0.4089 ± N/A
none 0 prompt_level_loose_acc ↑ 0.3216 ± 0.0201
none 0 prompt_level_strict_acc ↑ 0.2717 ± 0.0191
Prism ML org

You can check the appendix (B. Benchmark Evaluation Methodology) for more details on the eval setup:
All the eval numbers in the paper were ran through the same setup.
https://github.com/PrismML-Eng/Bonsai-demo/blob/main/1-bit-bonsai-8b-whitepaper.pdf

@diyangoodluck

Can you share your findings on the performance?

I was doing more prompt-based comparison to other similar-sized models in LM Studio and Bonsai-8B was lagging behind. I am not sure if the reason in backend of llama.cpp or there is Bonsai performance issue, so I'm curious what was your experience.

Sign up or log in to comment