Premature ending inside final answer after long thinking

#16
by perelmanych - opened

After thinking is done, model stops in the middle of the final response. I am using b10092 build of llama-server with official WebUI when model is loading the following is displayed, may be that is the root of the problem:

0.01.334.528 W load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
0.01.334.533 W load: special_eot_id is not in special_eog_ids - the tokenizer config may be incorrect

My cli command is:

llama-server ^
--model C:\Users\user.lmstudio\models\unsloth\Laguna-S-2.1-GGUF\Laguna-S-2.1-UD-Q8_K_XL-00001-of-00004.gguf ^
--alias Laguna-S-2.1 ^
--jinja ^
--chat-template-file laguna.jinja ^
--threads 9 ^
--threads-http 4 ^
--flash-attn on ^
--no-context-shift ^
--temp 1.0 --top-k 40 --top-p 1.0 --min-p 0.01 --repeat-penalty 1.0 --presence-penalty 2.0 ^
--ctx-size 65536 ^
--n-predict 65536 ^
--host 0.0.0.0 --port 8000 ^
--no-mmap ^
--n-gpu-layers 999 ^
--n-cpu-moe 44 ^
--reasoning on ^
--chat-template-kwargs "{"preserve_thinking":true}" ^
--batch_size 1024 --ubatch_size 1024

The only difference from the official jinja template is "\n" after thinking tag "" to make model always think.

Example of the answer:

...
Final Answer
\boxed{\frac{c \lambda^2}{2} + \frac{2 v^2}{(v + c \lambda)^2} - 1 > 0}
<//think>

To prove the inequality (\frac{c \lambda^2}{2} + \frac{2 v^2}{(v + c \lambda)^2} - 1 > 0) given (c > 0), (\lambda > 0), (c \lambda < v < 1), and ((v + c \lambda)^2 > 4c\

I put double slash in thinking end tag because with HF formatting the tag was disappearing, so in reality it is normal ending tag.

Sign up or log in to comment