0.00.146.849 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.146.852 I device_info: 0.00.249.664 I - CUDA0 : NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition (97244 MiB, 96677 MiB free) 0.00.249.678 I - CPU : AMD Ryzen Threadripper 9960X 24-Cores (257066 MiB, 257066 MiB free) 0.00.249.765 I system_info: n_threads = 24 (n_threads_batch = 24) / 48 | CUDA : ARCHS = 1200 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | BLACKWELL_NATIVE_FP4 = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | 0.00.249.850 I srv init: using 47 threads for HTTP server 0.00.250.054 I srv start: binding port with default address family 0.00.251.224 I srv llama_server: loading model 0.00.251.264 I srv load_model: loading model '/home/ripper/ornith-35b-lab/artifacts/quant/ornith-1.0-35b-IQ4_XS.gguf' 0.00.533.150 I srv load_model: [spec] estimated memory usage of draft model is 3176.45 MiB 0.00.533.186 I common_init_result: fitting params to device memory ... 0.00.533.187 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on) 0.02.569.364 W llama_context: n_ctx_seq (512) < n_ctx_train (262144) -- the full capacity of the model will not be utilized 0.02.584.209 I common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable) 0.02.616.143 I srv load_model: loading draft model '/home/ripper/ornith-35b-lab/artifacts/mtp/ornith-1.0-35b-mtp-chaincorr-h2r9fw125-h1r2-h3r6ce375-lr5e7-Q6_K.gguf' 0.02.957.174 W llama_context: n_ctx_seq (512) < n_ctx_train (262144) -- the full capacity of the model will not be utilized 0.02.966.075 I srv load_model: initializing slots, n_slots = 16 0.02.973.425 I common_context_can_seq_rm: the context supports bounded partial sequence removal 0.02.979.547 I common_speculative_impl_draft_mtp: adding speculative implementation 'draft-mtp' 0.02.979.550 I common_speculative_impl_draft_mtp: - n_max=1, n_min=0, p_min=0.00, n_embd=2048, backend_sampling=1 0.02.979.551 I common_speculative_impl_draft_mtp: - gpu_layers=99, cache_k=f16, cache_v=f16, ctx_tgt=yes, ctx_dft=yes, devices=[default] 0.02.979.679 I common_speculative_impl_draft_mtp: - draft sampler top_k=1 top_p=1.000 temp=1.000 0.02.979.705 I common_speculative_impl_draft_mtp: - fast backend MTP sampling enabled 0.02.979.713 I common_speculative_impl_draft_mtp: - backend draft sampler top_k=1 top_p=1.000 temp=1.000 contexts=1 0.02.998.515 I srv load_model: speculative decoding context initialized 0.02.998.518 I slot load_model: id 0 | task -1 | new slot, n_ctx = 512 0.02.998.527 I slot load_model: id 1 | task -1 | new slot, n_ctx = 512 0.02.998.527 I slot load_model: id 2 | task -1 | new slot, n_ctx = 512 0.02.998.527 I slot load_model: id 3 | task -1 | new slot, n_ctx = 512 0.02.998.528 I slot load_model: id 4 | task -1 | new slot, n_ctx = 512 0.02.998.528 I slot load_model: id 5 | task -1 | new slot, n_ctx = 512 0.02.998.528 I slot load_model: id 6 | task -1 | new slot, n_ctx = 512 0.02.998.528 I slot load_model: id 7 | task -1 | new slot, n_ctx = 512 0.02.998.528 I slot load_model: id 8 | task -1 | new slot, n_ctx = 512 0.02.998.528 I slot load_model: id 9 | task -1 | new slot, n_ctx = 512 0.02.998.529 I slot load_model: id 10 | task -1 | new slot, n_ctx = 512 0.02.998.529 I slot load_model: id 11 | task -1 | new slot, n_ctx = 512 0.02.998.529 I slot load_model: id 12 | task -1 | new slot, n_ctx = 512 0.02.998.529 I slot load_model: id 13 | task -1 | new slot, n_ctx = 512 0.02.998.529 I slot load_model: id 14 | task -1 | new slot, n_ctx = 512 0.02.998.530 I slot load_model: id 15 | task -1 | new slot, n_ctx = 512 0.02.998.578 I srv load_model: prompt cache is enabled, size limit: 8192 MiB 0.02.998.579 I srv load_model: use `--cache-ram 0` to disable the prompt cache 0.02.998.579 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391 0.02.998.579 I srv load_model: context checkpoints enabled, max = 32, min spacing = 8192 0.02.998.595 I srv init: idle slots will be saved to prompt cache upon starting a new task 0.03.007.350 I init: chat template, example_format: '<|im_start|>system You are a helpful assistant<|im_end|> <|im_start|>user Hello<|im_end|> <|im_start|>assistant Hi there<|im_end|> <|im_start|>user How are you?<|im_end|> <|im_start|>assistant ' 0.03.013.802 I srv init: init: chat template, thinking = 0 0.03.013.847 I srv llama_server: model loaded 0.03.013.852 I srv llama_server: server is listening on http://127.0.0.1:18240 0.03.013.858 I srv update_slots: all slots are idle 0.40.845.489 I srv operator(): Chat format: peg-native 0.40.845.494 I srv operator(): Chat format: peg-native 0.40.845.497 I srv operator(): Chat format: peg-native 0.40.845.502 I srv operator(): Chat format: peg-native 0.40.845.508 I srv operator(): Chat format: peg-native 0.40.845.515 I srv operator(): Chat format: peg-native 0.40.845.517 I srv operator(): Chat format: peg-native 0.40.845.544 I srv operator(): Chat format: peg-native 0.40.845.553 I srv operator(): Chat format: peg-native 0.40.845.634 I slot get_availabl: id 15 | task -1 | selected slot by LRU, t_last = -1 0.40.845.636 I srv get_availabl: updating prompt cache 0.40.845.641 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.645 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.646 I srv get_availabl: prompt cache update took 0.01 ms 0.40.845.684 I slot launch_slot_: id 15 | task 2 | processing task, is_child = 0 0.40.845.685 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.686 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.686 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.686 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.686 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.686 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.686 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.687 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.40.845.687 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 0.40.845.687 I slot process_sing: id 9 | task -1 | saving idle slot to prompt cache 0.40.845.687 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 0.40.845.687 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.40.845.687 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 0.40.845.687 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 0.40.845.688 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 0.40.845.688 I slot get_availabl: id 14 | task -1 | selected slot by LRU, t_last = -1 0.40.845.688 I srv get_availabl: updating prompt cache 0.40.845.690 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.691 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.691 I srv get_availabl: prompt cache update took 0.00 ms 0.40.845.702 I slot launch_slot_: id 14 | task 0 | processing task, is_child = 0 0.40.845.702 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.703 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.703 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.703 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.703 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.704 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.704 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.704 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.40.845.704 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 0.40.845.704 I slot process_sing: id 9 | task -1 | saving idle slot to prompt cache 0.40.845.705 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 0.40.845.705 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.40.845.705 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 0.40.845.705 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 0.40.845.706 I slot get_availabl: id 13 | task -1 | selected slot by LRU, t_last = -1 0.40.845.706 I srv get_availabl: updating prompt cache 0.40.845.707 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.707 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.708 I srv get_availabl: prompt cache update took 0.00 ms 0.40.845.719 I slot launch_slot_: id 13 | task 5 | processing task, is_child = 0 0.40.845.719 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.719 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.719 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.720 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.720 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.720 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.720 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.721 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.40.845.721 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 0.40.845.721 I slot process_sing: id 9 | task -1 | saving idle slot to prompt cache 0.40.845.722 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 0.40.845.722 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.40.845.722 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 0.40.845.723 I slot get_availabl: id 12 | task -1 | selected slot by LRU, t_last = -1 0.40.845.723 I srv get_availabl: updating prompt cache 0.40.845.724 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.726 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.726 I srv get_availabl: prompt cache update took 0.00 ms 0.40.845.737 I slot launch_slot_: id 12 | task 3 | processing task, is_child = 0 0.40.845.737 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.737 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.737 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.737 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.738 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.738 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.738 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.739 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.40.845.739 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 0.40.845.739 I slot process_sing: id 9 | task -1 | saving idle slot to prompt cache 0.40.845.739 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 0.40.845.739 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.40.845.740 I slot get_availabl: id 11 | task -1 | selected slot by LRU, t_last = -1 0.40.845.740 I srv get_availabl: updating prompt cache 0.40.845.740 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.740 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.741 I srv get_availabl: prompt cache update took 0.00 ms 0.40.845.752 I slot launch_slot_: id 11 | task 6 | processing task, is_child = 0 0.40.845.752 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.752 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.752 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.752 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.753 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.753 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.753 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.753 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.40.845.753 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 0.40.845.754 I slot process_sing: id 9 | task -1 | saving idle slot to prompt cache 0.40.845.754 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 0.40.845.754 I slot get_availabl: id 10 | task -1 | selected slot by LRU, t_last = -1 0.40.845.755 I srv get_availabl: updating prompt cache 0.40.845.755 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.755 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.757 I srv get_availabl: prompt cache update took 0.00 ms 0.40.845.766 I slot launch_slot_: id 10 | task 1 | processing task, is_child = 0 0.40.845.767 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.767 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.767 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.767 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.767 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.768 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.768 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.768 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.40.845.768 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 0.40.845.769 I slot process_sing: id 9 | task -1 | saving idle slot to prompt cache 0.40.845.769 I slot get_availabl: id 9 | task -1 | selected slot by LRU, t_last = -1 0.40.845.769 I srv get_availabl: updating prompt cache 0.40.845.770 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.771 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.771 I srv get_availabl: prompt cache update took 0.00 ms 0.40.845.780 I slot launch_slot_: id 9 | task 4 | processing task, is_child = 0 0.40.845.780 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.780 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.781 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.781 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.781 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.781 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.782 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.782 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.40.845.782 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 0.40.845.783 I slot get_availabl: id 8 | task -1 | selected slot by LRU, t_last = -1 0.40.845.783 I srv get_availabl: updating prompt cache 0.40.845.784 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.784 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.784 I srv get_availabl: prompt cache update took 0.00 ms 0.40.845.795 I slot launch_slot_: id 8 | task 8 | processing task, is_child = 0 0.40.845.796 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.797 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.797 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.797 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.797 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.798 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.798 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.798 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.40.845.799 I slot get_availabl: id 7 | task -1 | selected slot by LRU, t_last = -1 0.40.845.799 I srv get_availabl: updating prompt cache 0.40.845.799 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.40.845.801 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.40.845.801 I srv get_availabl: prompt cache update took 0.00 ms 0.40.845.809 I slot launch_slot_: id 7 | task 7 | processing task, is_child = 0 0.40.845.810 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.40.845.810 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.40.845.810 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.40.845.810 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.40.845.810 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.40.845.811 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.40.845.811 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.40.845.846 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.845.858 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.845.861 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.845.865 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.845.871 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.845.875 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.845.883 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.845.890 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.846.103 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.40.846.295 I srv operator(): Chat format: peg-native 0.40.848.113 I srv operator(): Chat format: peg-native 0.40.849.418 I srv operator(): Chat format: peg-native 0.40.849.767 I srv operator(): Chat format: peg-native 0.40.850.773 I srv operator(): Chat format: peg-native 0.40.850.864 I srv operator(): Chat format: peg-native 0.40.852.523 I srv operator(): Chat format: peg-native 0.41.095.382 I slot get_availabl: id 6 | task -1 | selected slot by LRU, t_last = -1 0.41.095.386 I srv get_availabl: updating prompt cache 0.41.095.391 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.41.095.395 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.41.095.395 I srv get_availabl: prompt cache update took 0.01 ms 0.41.095.468 I slot launch_slot_: id 6 | task 10 | processing task, is_child = 0 0.41.095.469 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.41.095.470 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.41.095.470 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.41.095.470 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.41.095.470 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.41.095.470 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.41.095.471 I slot get_availabl: id 5 | task -1 | selected slot by LRU, t_last = -1 0.41.095.472 I srv get_availabl: updating prompt cache 0.41.095.472 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.41.095.472 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.41.095.473 I srv get_availabl: prompt cache update took 0.00 ms 0.41.095.487 I slot launch_slot_: id 5 | task 11 | processing task, is_child = 0 0.41.095.487 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.41.095.487 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.41.095.488 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.41.095.488 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.41.095.488 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.41.095.490 I slot get_availabl: id 4 | task -1 | selected slot by LRU, t_last = -1 0.41.095.490 I srv get_availabl: updating prompt cache 0.41.095.491 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.41.095.491 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.41.095.491 I srv get_availabl: prompt cache update took 0.00 ms 0.41.095.503 I slot launch_slot_: id 4 | task 12 | processing task, is_child = 0 0.41.095.505 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.41.095.505 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.41.095.505 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.41.095.505 I slot process_sing: id 3 | task -1 | saving idle slot to prompt cache 0.41.095.507 I slot get_availabl: id 3 | task -1 | selected slot by LRU, t_last = -1 0.41.095.507 I srv get_availabl: updating prompt cache 0.41.095.507 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.41.095.508 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.41.095.509 I srv get_availabl: prompt cache update took 0.00 ms 0.41.095.519 I slot launch_slot_: id 3 | task 13 | processing task, is_child = 0 0.41.095.520 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.41.095.521 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.41.095.521 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.41.095.523 I slot get_availabl: id 2 | task -1 | selected slot by LRU, t_last = -1 0.41.095.523 I srv get_availabl: updating prompt cache 0.41.095.523 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.41.095.523 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.41.095.524 I srv get_availabl: prompt cache update took 0.00 ms 0.41.095.536 I slot launch_slot_: id 2 | task 14 | processing task, is_child = 0 0.41.095.537 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.41.095.538 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.41.095.539 I slot get_availabl: id 1 | task -1 | selected slot by LRU, t_last = -1 0.41.095.539 I srv get_availabl: updating prompt cache 0.41.095.540 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.41.095.540 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.41.095.540 I srv get_availabl: prompt cache update took 0.00 ms 0.41.095.551 I slot launch_slot_: id 1 | task 15 | processing task, is_child = 0 0.41.095.553 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 0.41.095.554 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1 0.41.095.554 I srv get_availabl: updating prompt cache 0.41.095.554 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000 0.41.095.555 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 8192 tokens, 8589934592 est) 0.41.095.555 I srv get_availabl: prompt cache update took 0.00 ms 0.41.095.566 I slot launch_slot_: id 0 | task 16 | processing task, is_child = 0 0.41.095.677 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.41.095.683 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.41.095.697 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.41.095.896 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.41.095.900 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.41.095.904 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.41.095.918 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.41.109.822 I slot create_check: id 7 | task 7 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.41.123.512 I slot create_check: id 8 | task 8 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.41.137.138 I slot create_check: id 9 | task 4 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.41.150.832 I slot create_check: id 10 | task 1 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.41.164.351 I slot create_check: id 11 | task 6 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.41.177.857 I slot create_check: id 12 | task 3 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.41.191.714 I slot create_check: id 13 | task 5 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.41.205.402 I slot create_check: id 14 | task 0 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.41.219.310 I slot create_check: id 15 | task 2 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.41.433.403 I slot create_check: id 0 | task 16 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.41.446.948 I slot create_check: id 1 | task 15 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.41.460.770 I slot create_check: id 2 | task 14 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.41.474.347 I slot create_check: id 3 | task 13 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.41.487.871 I slot create_check: id 4 | task 12 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.41.501.356 I slot create_check: id 5 | task 11 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.41.514.779 I slot create_check: id 6 | task 10 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.46.418.396 I slot print_timing: id 7 | task 7 | n_decoded = 101, tg = 20.17 t/s, tg_3s = 20.17 t/s 0.46.420.109 I slot print_timing: id 12 | task 3 | n_decoded = 101, tg = 20.17 t/s, tg_3s = 20.17 t/s 0.46.421.418 I slot print_timing: id 15 | task 2 | n_decoded = 101, tg = 20.18 t/s, tg_3s = 20.18 t/s 0.46.511.573 I slot print_timing: id 1 | task 15 | n_decoded = 101, tg = 20.69 t/s, tg_3s = 20.69 t/s 0.46.512.866 I slot print_timing: id 4 | task 12 | n_decoded = 101, tg = 20.69 t/s, tg_3s = 20.69 t/s 0.46.513.296 I slot print_timing: id 5 | task 11 | n_decoded = 101, tg = 20.69 t/s, tg_3s = 20.69 t/s 0.46.801.643 I slot print_timing: id 11 | task 6 | n_decoded = 100, tg = 18.56 t/s, tg_3s = 18.56 t/s 0.46.802.565 I slot print_timing: id 13 | task 5 | n_decoded = 100, tg = 18.56 t/s, tg_3s = 18.56 t/s 0.46.892.376 I slot print_timing: id 0 | task 16 | n_decoded = 100, tg = 19.00 t/s, tg_3s = 19.00 t/s 0.46.893.285 I slot print_timing: id 2 | task 14 | n_decoded = 100, tg = 19.00 t/s, tg_3s = 19.00 t/s 0.46.894.847 I slot print_timing: id 6 | task 10 | n_decoded = 100, tg = 19.01 t/s, tg_3s = 19.01 t/s 0.47.084.774 I slot print_timing: id 7 | task 7 | prompt eval time = 564.59 ms / 25 tokens ( 22.58 ms per token, 44.28 tokens per second) 0.47.084.777 I slot print_timing: id 7 | task 7 | eval time = 5674.34 ms / 114 tokens ( 49.77 ms per token, 20.09 tokens per second) 0.47.084.778 I slot print_timing: id 7 | task 7 | total time = 6238.93 ms / 139 tokens 0.47.084.783 I slot print_timing: id 7 | task 7 | graphs reused = 1 0.47.084.786 I slot print_timing: id 7 | task 7 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.47.084.814 I statistics draft-mtp: #calls(b,g,a) = 16 58 913, #gen drafts = 921, #acc drafts = 674, #gen tokens = 921, #acc tokens = 674, #mean acc len = 1.74, #acc rate/pos = (0.738), dur(b,g,a) = 0.014, 84.136, 1.077 ms 0.47.084.855 I slot release: id 7 | task 7 | stop processing: n_tokens = 139, truncated = 0 0.47.087.044 I slot print_timing: id 12 | task 3 | prompt eval time = 567.85 ms / 25 tokens ( 22.71 ms per token, 44.03 tokens per second) 0.47.087.045 I slot print_timing: id 12 | task 3 | eval time = 5673.31 ms / 114 tokens ( 49.77 ms per token, 20.09 tokens per second) 0.47.087.046 I slot print_timing: id 12 | task 3 | total time = 6241.16 ms / 139 tokens 0.47.087.047 I slot print_timing: id 12 | task 3 | graphs reused = 1 0.47.087.049 I slot print_timing: id 12 | task 3 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.47.087.056 I statistics draft-mtp: #calls(b,g,a) = 16 58 918, #gen drafts = 921, #acc drafts = 679, #gen tokens = 921, #acc tokens = 679, #mean acc len = 1.74, #acc rate/pos = (0.740), dur(b,g,a) = 0.014, 84.136, 1.082 ms 0.47.087.078 I slot release: id 12 | task 3 | stop processing: n_tokens = 139, truncated = 0 0.47.088.233 I slot print_timing: id 15 | task 2 | prompt eval time = 569.90 ms / 25 tokens ( 22.80 ms per token, 43.87 tokens per second) 0.47.088.235 I slot print_timing: id 15 | task 2 | eval time = 5672.43 ms / 114 tokens ( 49.76 ms per token, 20.10 tokens per second) 0.47.088.235 I slot print_timing: id 15 | task 2 | total time = 6242.33 ms / 139 tokens 0.47.088.236 I slot print_timing: id 15 | task 2 | graphs reused = 1 0.47.088.237 I slot print_timing: id 15 | task 2 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.47.088.244 I statistics draft-mtp: #calls(b,g,a) = 16 58 921, #gen drafts = 921, #acc drafts = 681, #gen tokens = 921, #acc tokens = 681, #mean acc len = 1.74, #acc rate/pos = (0.739), dur(b,g,a) = 0.014, 84.136, 1.086 ms 0.47.088.264 I slot release: id 15 | task 2 | stop processing: n_tokens = 139, truncated = 0 0.47.093.701 I srv operator(): Chat format: peg-native 0.47.096.202 I srv operator(): Chat format: peg-native 0.47.096.220 I srv operator(): Chat format: peg-native 0.47.171.885 I slot print_timing: id 1 | task 15 | prompt eval time = 534.85 ms / 25 tokens ( 21.39 ms per token, 46.74 tokens per second) 0.47.171.888 I slot print_timing: id 1 | task 15 | eval time = 5541.40 ms / 114 tokens ( 48.61 ms per token, 20.57 tokens per second) 0.47.171.889 I slot print_timing: id 1 | task 15 | total time = 6076.25 ms / 139 tokens 0.47.171.890 I slot print_timing: id 1 | task 15 | graphs reused = 1 0.47.171.892 I slot print_timing: id 1 | task 15 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.47.171.909 I statistics draft-mtp: #calls(b,g,a) = 16 59 923, #gen drafts = 934, #acc drafts = 683, #gen tokens = 934, #acc tokens = 683, #mean acc len = 1.74, #acc rate/pos = (0.740), dur(b,g,a) = 0.014, 88.464, 1.090 ms 0.47.171.933 I slot release: id 1 | task 15 | stop processing: n_tokens = 139, truncated = 0 0.47.173.042 I slot print_timing: id 4 | task 12 | prompt eval time = 536.55 ms / 25 tokens ( 21.46 ms per token, 46.59 tokens per second) 0.47.173.044 I slot print_timing: id 4 | task 12 | eval time = 5540.79 ms / 114 tokens ( 48.60 ms per token, 20.57 tokens per second) 0.47.173.044 I slot print_timing: id 4 | task 12 | total time = 6077.33 ms / 139 tokens 0.47.173.045 I slot print_timing: id 4 | task 12 | graphs reused = 1 0.47.173.046 I slot print_timing: id 4 | task 12 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.47.173.054 I statistics draft-mtp: #calls(b,g,a) = 16 59 926, #gen drafts = 934, #acc drafts = 685, #gen tokens = 934, #acc tokens = 685, #mean acc len = 1.74, #acc rate/pos = (0.740), dur(b,g,a) = 0.014, 88.464, 1.093 ms 0.47.173.074 I slot release: id 4 | task 12 | stop processing: n_tokens = 139, truncated = 0 0.47.173.500 I slot print_timing: id 5 | task 11 | prompt eval time = 537.09 ms / 25 tokens ( 21.48 ms per token, 46.55 tokens per second) 0.47.173.502 I slot print_timing: id 5 | task 11 | eval time = 5540.68 ms / 114 tokens ( 48.60 ms per token, 20.58 tokens per second) 0.47.173.503 I slot print_timing: id 5 | task 11 | total time = 6077.77 ms / 139 tokens 0.47.173.503 I slot print_timing: id 5 | task 11 | graphs reused = 1 0.47.173.505 I slot print_timing: id 5 | task 11 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.47.173.511 I statistics draft-mtp: #calls(b,g,a) = 16 59 927, #gen drafts = 934, #acc drafts = 686, #gen tokens = 934, #acc tokens = 686, #mean acc len = 1.74, #acc rate/pos = (0.740), dur(b,g,a) = 0.014, 88.464, 1.095 ms 0.47.173.528 I slot release: id 5 | task 11 | stop processing: n_tokens = 139, truncated = 0 0.47.176.421 I slot get_availabl: id 1 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 0.47.176.424 I srv get_availabl: updating prompt cache 0.47.176.472 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.180.329 I srv operator(): Chat format: peg-native 0.47.182.251 I srv operator(): Chat format: peg-native 0.47.182.725 I srv operator(): Chat format: peg-native 0.47.204.230 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 0.47.204.235 I srv update: - cache state: 1 prompts, 129.598 MiB (limits: 8192.000 MiB, 8192 tokens, 8786 est) 0.47.204.236 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.47.204.236 I srv get_availabl: prompt cache update took 27.81 ms 0.47.204.317 I slot launch_slot_: id 1 | task 77 | processing task, is_child = 0 0.47.204.318 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 0.47.204.350 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.351 I srv alloc: - prompt is already in the cache, skipping 0.47.204.352 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.47.204.368 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.370 I srv alloc: - prompt is already in the cache, skipping 0.47.204.370 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.47.204.387 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.388 I srv alloc: - prompt is already in the cache, skipping 0.47.204.389 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 0.47.204.409 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.410 I srv alloc: - prompt is already in the cache, skipping 0.47.204.411 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 0.47.204.426 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.428 I srv alloc: - prompt is already in the cache, skipping 0.47.204.431 I slot get_availabl: id 4 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 0.47.204.432 I srv get_availabl: updating prompt cache 0.47.204.447 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.449 I srv alloc: - prompt is already in the cache, skipping 0.47.204.449 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 0.47.204.450 I srv update: - cache state: 1 prompts, 129.598 MiB (limits: 8192.000 MiB, 8192 tokens, 8786 est) 0.47.204.451 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.47.204.452 I srv get_availabl: prompt cache update took 0.02 ms 0.47.204.479 I slot launch_slot_: id 4 | task 79 | processing task, is_child = 0 0.47.204.481 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 0.47.204.502 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.504 I srv alloc: - prompt is already in the cache, skipping 0.47.204.504 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.47.204.525 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.527 I srv alloc: - prompt is already in the cache, skipping 0.47.204.527 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 0.47.204.549 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.550 I srv alloc: - prompt is already in the cache, skipping 0.47.204.551 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 0.47.204.571 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.572 I srv alloc: - prompt is already in the cache, skipping 0.47.204.575 I slot get_availabl: id 5 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.022 0.47.204.576 I srv get_availabl: updating prompt cache 0.47.204.597 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.599 I srv alloc: - prompt is already in the cache, skipping 0.47.204.599 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.107 0.47.204.600 I srv update: - cache state: 1 prompts, 129.598 MiB (limits: 8192.000 MiB, 8192 tokens, 8786 est) 0.47.204.602 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.47.204.602 I srv get_availabl: prompt cache update took 0.03 ms 0.47.204.624 I slot launch_slot_: id 5 | task 78 | processing task, is_child = 0 0.47.204.626 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.47.204.648 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.649 I srv alloc: - prompt is already in the cache, skipping 0.47.204.649 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 0.47.204.671 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.673 I srv alloc: - prompt is already in the cache, skipping 0.47.204.673 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 0.47.204.695 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.696 I srv alloc: - prompt is already in the cache, skipping 0.47.204.699 I slot get_availabl: id 7 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 0.47.204.699 I srv get_availabl: updating prompt cache 0.47.204.721 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.723 I srv alloc: - prompt is already in the cache, skipping 0.47.204.723 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 0.47.204.724 I srv update: - cache state: 1 prompts, 129.598 MiB (limits: 8192.000 MiB, 8192 tokens, 8786 est) 0.47.204.726 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.47.204.726 I srv get_availabl: prompt cache update took 0.03 ms 0.47.204.745 I slot launch_slot_: id 7 | task 80 | processing task, is_child = 0 0.47.204.746 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 0.47.204.767 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.769 I srv alloc: - prompt is already in the cache, skipping 0.47.204.769 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 0.47.204.790 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.792 I srv alloc: - prompt is already in the cache, skipping 0.47.204.794 I slot get_availabl: id 12 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.022 0.47.204.795 I srv get_availabl: updating prompt cache 0.47.204.816 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.817 I srv alloc: - prompt is already in the cache, skipping 0.47.204.818 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.107 0.47.204.819 I srv update: - cache state: 1 prompts, 129.598 MiB (limits: 8192.000 MiB, 8192 tokens, 8786 est) 0.47.204.821 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.47.204.821 I srv get_availabl: prompt cache update took 0.03 ms 0.47.204.839 I slot launch_slot_: id 12 | task 81 | processing task, is_child = 0 0.47.204.840 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 0.47.204.861 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.862 I srv alloc: - prompt is already in the cache, skipping 0.47.204.865 I slot get_availabl: id 15 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 0.47.204.866 I srv get_availabl: updating prompt cache 0.47.204.887 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.47.204.888 I srv alloc: - prompt is already in the cache, skipping 0.47.204.889 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 0.47.204.890 I srv update: - cache state: 1 prompts, 129.598 MiB (limits: 8192.000 MiB, 8192 tokens, 8786 est) 0.47.204.891 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.47.204.891 I srv get_availabl: prompt cache update took 0.03 ms 0.47.204.909 I slot launch_slot_: id 15 | task 82 | processing task, is_child = 0 0.47.208.144 I slot operator(): id 1 | task 77 | Checking checkpoint with [20, 20] against 3... 0.47.208.146 W slot operator(): id 1 | task 77 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.47.208.149 W slot operator(): id 1 | task 77 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 0.47.209.735 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.47.209.756 I slot operator(): id 4 | task 79 | Checking checkpoint with [20, 20] against 24... 0.47.213.374 W slot operator(): id 4 | task 79 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 0.47.213.421 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.47.213.429 I slot operator(): id 5 | task 78 | Checking checkpoint with [20, 20] against 3... 0.47.213.431 W slot operator(): id 5 | task 78 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.47.213.431 W slot operator(): id 5 | task 78 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 0.47.215.025 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.47.215.042 I slot operator(): id 7 | task 80 | Checking checkpoint with [20, 20] against 3... 0.47.215.044 W slot operator(): id 7 | task 80 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.47.215.046 W slot operator(): id 7 | task 80 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 0.47.217.032 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.47.217.055 I slot operator(): id 12 | task 81 | Checking checkpoint with [20, 20] against 3... 0.47.217.057 W slot operator(): id 12 | task 81 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.47.217.059 W slot operator(): id 12 | task 81 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 0.47.219.216 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.47.219.230 I slot operator(): id 15 | task 82 | Checking checkpoint with [20, 20] against 24... 0.47.222.792 W slot operator(): id 15 | task 82 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 0.47.222.813 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.47.388.698 I slot create_check: id 1 | task 77 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.47.402.268 I slot create_check: id 5 | task 78 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.47.415.901 I slot create_check: id 7 | task 80 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.47.429.500 I slot create_check: id 12 | task 81 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.48.290.339 I slot print_timing: id 8 | task 8 | n_decoded = 101, tg = 14.68 t/s, tg_3s = 14.68 t/s 0.48.290.776 I slot print_timing: id 9 | task 4 | n_decoded = 101, tg = 14.68 t/s, tg_3s = 14.68 t/s 0.48.291.211 I slot print_timing: id 10 | task 1 | n_decoded = 101, tg = 14.68 t/s, tg_3s = 14.68 t/s 0.48.387.160 I slot print_timing: id 14 | task 0 | n_decoded = 101, tg = 14.49 t/s, tg_3s = 14.49 t/s 0.48.477.624 I slot print_timing: id 3 | task 13 | n_decoded = 101, tg = 14.75 t/s, tg_3s = 14.75 t/s 0.48.666.121 I slot print_timing: id 11 | task 6 | prompt eval time = 567.15 ms / 28 tokens ( 20.26 ms per token, 49.37 tokens per second) 0.48.666.124 I slot print_timing: id 11 | task 6 | eval time = 7253.08 ms / 128 tokens ( 56.66 ms per token, 17.65 tokens per second) 0.48.666.124 I slot print_timing: id 11 | task 6 | total time = 7820.23 ms / 156 tokens 0.48.666.125 I slot print_timing: id 11 | task 6 | graphs reused = 1 0.48.666.127 I slot print_timing: id 11 | task 6 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.48.666.140 I statistics draft-mtp: #calls(b,g,a) = 22 73 1132, #gen drafts = 1146, #acc drafts = 816, #gen tokens = 1146, #acc tokens = 816, #mean acc len = 1.72, #acc rate/pos = (0.721), dur(b,g,a) = 0.016, 113.551, 1.343 ms 0.48.666.164 I slot release: id 11 | task 6 | stop processing: n_tokens = 155, truncated = 0 0.48.666.366 I slot print_timing: id 13 | task 5 | prompt eval time = 568.61 ms / 28 tokens ( 20.31 ms per token, 49.24 tokens per second) 0.48.666.368 I slot print_timing: id 13 | task 5 | eval time = 7251.87 ms / 128 tokens ( 56.66 ms per token, 17.65 tokens per second) 0.48.666.368 I slot print_timing: id 13 | task 5 | total time = 7820.48 ms / 156 tokens 0.48.666.368 I slot print_timing: id 13 | task 5 | graphs reused = 1 0.48.666.370 I slot print_timing: id 13 | task 5 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.48.666.371 I statistics draft-mtp: #calls(b,g,a) = 22 73 1132, #gen drafts = 1146, #acc drafts = 816, #gen tokens = 1146, #acc tokens = 816, #mean acc len = 1.72, #acc rate/pos = (0.721), dur(b,g,a) = 0.016, 113.551, 1.343 ms 0.48.666.384 I slot release: id 13 | task 5 | stop processing: n_tokens = 155, truncated = 0 0.48.676.095 I srv operator(): Chat format: peg-native 0.48.676.571 I srv operator(): Chat format: peg-native 0.48.748.542 I slot print_timing: id 0 | task 16 | prompt eval time = 534.27 ms / 28 tokens ( 19.08 ms per token, 52.41 tokens per second) 0.48.748.545 I slot print_timing: id 0 | task 16 | eval time = 7118.68 ms / 128 tokens ( 55.61 ms per token, 17.98 tokens per second) 0.48.748.545 I slot print_timing: id 0 | task 16 | total time = 7652.94 ms / 156 tokens 0.48.748.546 I slot print_timing: id 0 | task 16 | graphs reused = 1 0.48.748.548 I slot print_timing: id 0 | task 16 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.48.748.559 I statistics draft-mtp: #calls(b,g,a) = 22 74 1146, #gen drafts = 1157, #acc drafts = 823, #gen tokens = 1157, #acc tokens = 823, #mean acc len = 1.72, #acc rate/pos = (0.718), dur(b,g,a) = 0.016, 116.641, 1.362 ms 0.48.748.581 I slot release: id 0 | task 16 | stop processing: n_tokens = 155, truncated = 0 0.48.748.783 I slot print_timing: id 2 | task 14 | prompt eval time = 535.41 ms / 28 tokens ( 19.12 ms per token, 52.30 tokens per second) 0.48.748.784 I slot print_timing: id 2 | task 14 | eval time = 7117.72 ms / 128 tokens ( 55.61 ms per token, 17.98 tokens per second) 0.48.748.785 I slot print_timing: id 2 | task 14 | total time = 7653.13 ms / 156 tokens 0.48.748.785 I slot print_timing: id 2 | task 14 | graphs reused = 1 0.48.748.786 I slot print_timing: id 2 | task 14 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.48.748.788 I statistics draft-mtp: #calls(b,g,a) = 22 74 1146, #gen drafts = 1157, #acc drafts = 823, #gen tokens = 1157, #acc tokens = 823, #mean acc len = 1.72, #acc rate/pos = (0.718), dur(b,g,a) = 0.016, 116.641, 1.362 ms 0.48.748.803 I slot release: id 2 | task 14 | stop processing: n_tokens = 155, truncated = 0 0.48.749.003 I slot print_timing: id 6 | task 10 | prompt eval time = 537.54 ms / 28 tokens ( 19.20 ms per token, 52.09 tokens per second) 0.48.749.005 I slot print_timing: id 6 | task 10 | eval time = 7115.56 ms / 128 tokens ( 55.59 ms per token, 17.99 tokens per second) 0.48.749.006 I slot print_timing: id 6 | task 10 | total time = 7653.10 ms / 156 tokens 0.48.749.006 I slot print_timing: id 6 | task 10 | graphs reused = 1 0.48.749.006 I slot print_timing: id 6 | task 10 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.48.749.007 I statistics draft-mtp: #calls(b,g,a) = 22 74 1146, #gen drafts = 1157, #acc drafts = 823, #gen tokens = 1157, #acc tokens = 823, #mean acc len = 1.72, #acc rate/pos = (0.718), dur(b,g,a) = 0.016, 116.641, 1.362 ms 0.48.749.022 I slot release: id 6 | task 10 | stop processing: n_tokens = 155, truncated = 0 0.48.752.984 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 0.48.752.985 I srv get_availabl: updating prompt cache 0.48.753.026 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.758.998 I srv operator(): Chat format: peg-native 0.48.759.116 I srv operator(): Chat format: peg-native 0.48.759.243 I srv operator(): Chat format: peg-native 0.48.781.060 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 0.48.781.065 I srv update: - cache state: 2 prompts, 259.657 MiB (limits: 8192.000 MiB, 8192 tokens, 9275 est) 0.48.781.066 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.48.781.067 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.48.781.068 I srv get_availabl: prompt cache update took 28.08 ms 0.48.781.150 I slot launch_slot_: id 0 | task 98 | processing task, is_child = 0 0.48.781.152 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 0.48.781.171 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.173 I srv alloc: - prompt is already in the cache, skipping 0.48.781.174 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.48.781.191 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.192 I srv alloc: - prompt is already in the cache, skipping 0.48.781.193 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.48.781.209 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.210 I srv alloc: - prompt is already in the cache, skipping 0.48.781.211 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 0.48.781.226 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.227 I srv alloc: - prompt is already in the cache, skipping 0.48.781.232 I slot get_availabl: id 2 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 0.48.781.233 I srv get_availabl: updating prompt cache 0.48.781.247 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.248 I srv alloc: - prompt is already in the cache, skipping 0.48.781.248 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 0.48.781.249 I srv update: - cache state: 2 prompts, 259.657 MiB (limits: 8192.000 MiB, 8192 tokens, 9275 est) 0.48.781.249 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.48.781.251 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.48.781.251 I srv get_availabl: prompt cache update took 0.02 ms 0.48.781.268 I slot launch_slot_: id 2 | task 99 | processing task, is_child = 0 0.48.781.269 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.48.781.284 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.285 I srv alloc: - prompt is already in the cache, skipping 0.48.781.285 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.48.781.299 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.300 I srv alloc: - prompt is already in the cache, skipping 0.48.781.301 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 0.48.781.314 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.316 I srv alloc: - prompt is already in the cache, skipping 0.48.781.317 I slot get_availabl: id 6 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 0.48.781.319 I srv get_availabl: updating prompt cache 0.48.781.333 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.334 I srv alloc: - prompt is already in the cache, skipping 0.48.781.334 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 0.48.781.335 I srv update: - cache state: 2 prompts, 259.657 MiB (limits: 8192.000 MiB, 8192 tokens, 9275 est) 0.48.781.335 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.48.781.336 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.48.781.336 I srv get_availabl: prompt cache update took 0.02 ms 0.48.781.352 I slot launch_slot_: id 6 | task 100 | processing task, is_child = 0 0.48.781.353 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.48.781.367 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.368 I srv alloc: - prompt is already in the cache, skipping 0.48.781.369 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 0.48.781.383 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.384 I srv alloc: - prompt is already in the cache, skipping 0.48.781.386 I slot get_availabl: id 11 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 0.48.781.386 I srv get_availabl: updating prompt cache 0.48.781.405 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.406 I srv alloc: - prompt is already in the cache, skipping 0.48.781.406 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 0.48.781.407 I srv update: - cache state: 2 prompts, 259.657 MiB (limits: 8192.000 MiB, 8192 tokens, 9275 est) 0.48.781.407 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.48.781.407 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.48.781.409 I srv get_availabl: prompt cache update took 0.02 ms 0.48.781.422 I slot launch_slot_: id 11 | task 101 | processing task, is_child = 0 0.48.781.424 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 0.48.781.437 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.438 I srv alloc: - prompt is already in the cache, skipping 0.48.781.440 I slot get_availabl: id 13 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 0.48.781.441 I srv get_availabl: updating prompt cache 0.48.781.456 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.48.781.457 I srv alloc: - prompt is already in the cache, skipping 0.48.781.457 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 0.48.781.458 I srv update: - cache state: 2 prompts, 259.657 MiB (limits: 8192.000 MiB, 8192 tokens, 9275 est) 0.48.781.458 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.48.781.458 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.48.781.458 I srv get_availabl: prompt cache update took 0.02 ms 0.48.781.473 I slot launch_slot_: id 13 | task 102 | processing task, is_child = 0 0.48.784.637 I slot operator(): id 0 | task 98 | Checking checkpoint with [23, 23] against 27... 0.48.788.155 W slot operator(): id 0 | task 98 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 0.48.788.205 I slot operator(): id 2 | task 99 | Checking checkpoint with [23, 23] against 3... 0.48.788.209 W slot operator(): id 2 | task 99 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.48.788.210 W slot operator(): id 2 | task 99 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 0.48.788.211 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.48.789.931 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.48.789.962 I slot operator(): id 6 | task 100 | Checking checkpoint with [23, 23] against 3... 0.48.789.966 W slot operator(): id 6 | task 100 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.48.789.968 W slot operator(): id 6 | task 100 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 0.48.791.452 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.48.791.466 I slot operator(): id 11 | task 101 | Checking checkpoint with [23, 23] against 27... 0.48.794.983 W slot operator(): id 11 | task 101 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 0.48.795.008 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.48.795.032 I slot operator(): id 13 | task 102 | Checking checkpoint with [23, 23] against 3... 0.48.795.033 W slot operator(): id 13 | task 102 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.48.795.034 W slot operator(): id 13 | task 102 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 0.48.796.588 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.48.948.377 I slot create_check: id 2 | task 99 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.48.961.558 I slot create_check: id 6 | task 100 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.48.975.047 I slot create_check: id 13 | task 102 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.50.304.670 I slot print_timing: id 8 | task 8 | prompt eval time = 565.25 ms / 29 tokens ( 19.49 ms per token, 51.30 tokens per second) 0.50.304.675 I slot print_timing: id 8 | task 8 | eval time = 8893.54 ms / 128 tokens ( 69.48 ms per token, 14.39 tokens per second) 0.50.304.676 I slot print_timing: id 8 | task 8 | total time = 9458.79 ms / 157 tokens 0.50.304.676 I slot print_timing: id 8 | task 8 | graphs reused = 1 0.50.304.679 I slot print_timing: id 8 | task 8 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 0.50.304.694 I statistics draft-mtp: #calls(b,g,a) = 27 89 1373, #gen drafts = 1386, #acc drafts = 971, #gen tokens = 1386, #acc tokens = 971, #mean acc len = 1.71, #acc rate/pos = (0.707), dur(b,g,a) = 0.018, 142.954, 1.624 ms 0.50.304.718 I slot release: id 8 | task 8 | stop processing: n_tokens = 156, truncated = 0 0.50.304.900 I slot print_timing: id 9 | task 4 | prompt eval time = 565.87 ms / 29 tokens ( 19.51 ms per token, 51.25 tokens per second) 0.50.304.902 I slot print_timing: id 9 | task 4 | eval time = 8893.17 ms / 128 tokens ( 69.48 ms per token, 14.39 tokens per second) 0.50.304.902 I slot print_timing: id 9 | task 4 | total time = 9459.03 ms / 157 tokens 0.50.304.903 I slot print_timing: id 9 | task 4 | graphs reused = 1 0.50.304.903 I slot print_timing: id 9 | task 4 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 0.50.304.905 I statistics draft-mtp: #calls(b,g,a) = 27 89 1373, #gen drafts = 1386, #acc drafts = 971, #gen tokens = 1386, #acc tokens = 971, #mean acc len = 1.71, #acc rate/pos = (0.707), dur(b,g,a) = 0.018, 142.954, 1.624 ms 0.50.304.926 I slot release: id 9 | task 4 | stop processing: n_tokens = 156, truncated = 0 0.50.305.105 I slot print_timing: id 10 | task 1 | prompt eval time = 566.53 ms / 29 tokens ( 19.54 ms per token, 51.19 tokens per second) 0.50.305.107 I slot print_timing: id 10 | task 1 | eval time = 8892.70 ms / 128 tokens ( 69.47 ms per token, 14.39 tokens per second) 0.50.305.108 I slot print_timing: id 10 | task 1 | total time = 9459.23 ms / 157 tokens 0.50.305.108 I slot print_timing: id 10 | task 1 | graphs reused = 1 0.50.305.109 I slot print_timing: id 10 | task 1 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 0.50.305.110 I statistics draft-mtp: #calls(b,g,a) = 27 89 1373, #gen drafts = 1386, #acc drafts = 971, #gen tokens = 1386, #acc tokens = 971, #mean acc len = 1.71, #acc rate/pos = (0.707), dur(b,g,a) = 0.018, 142.954, 1.624 ms 0.50.305.129 I slot release: id 10 | task 1 | stop processing: n_tokens = 156, truncated = 0 0.50.315.610 I srv operator(): Chat format: peg-native 0.50.315.711 I srv operator(): Chat format: peg-native 0.50.315.731 I srv operator(): Chat format: peg-native 0.50.384.496 I slot print_timing: id 14 | task 0 | prompt eval time = 569.26 ms / 29 tokens ( 19.63 ms per token, 50.94 tokens per second) 0.50.384.499 I slot print_timing: id 14 | task 0 | eval time = 8969.32 ms / 128 tokens ( 70.07 ms per token, 14.27 tokens per second) 0.50.384.499 I slot print_timing: id 14 | task 0 | total time = 9538.59 ms / 157 tokens 0.50.384.500 I slot print_timing: id 14 | task 0 | graphs reused = 1 0.50.384.502 I slot print_timing: id 14 | task 0 | draft acceptance = 0.41573 ( 37 accepted / 89 generated), mean acceptance length = 1.42, acceptance rate per position = (0.416) 0.50.384.515 I statistics draft-mtp: #calls(b,g,a) = 27 90 1386, #gen drafts = 1398, #acc drafts = 981, #gen tokens = 1398, #acc tokens = 981, #mean acc len = 1.71, #acc rate/pos = (0.708), dur(b,g,a) = 0.018, 145.386, 1.641 ms 0.50.384.538 I slot release: id 14 | task 0 | stop processing: n_tokens = 156, truncated = 0 0.50.389.437 I slot get_availabl: id 8 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 0.50.389.439 I srv get_availabl: updating prompt cache 0.50.389.480 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.393.797 I srv operator(): Chat format: peg-native 0.50.417.162 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 0.50.417.168 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.50.417.168 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.50.417.169 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.50.417.169 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.50.417.170 I srv get_availabl: prompt cache update took 27.73 ms 0.50.417.250 I slot launch_slot_: id 8 | task 119 | processing task, is_child = 0 0.50.417.254 I slot process_sing: id 9 | task -1 | saving idle slot to prompt cache 0.50.417.274 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.276 I srv alloc: - prompt is already in the cache, skipping 0.50.417.277 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 0.50.417.293 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.294 I srv alloc: - prompt is already in the cache, skipping 0.50.417.295 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 0.50.417.311 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.313 I srv alloc: - prompt is already in the cache, skipping 0.50.417.316 I slot get_availabl: id 9 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 0.50.417.316 I srv get_availabl: updating prompt cache 0.50.417.331 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.332 I srv alloc: - prompt is already in the cache, skipping 0.50.417.333 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 0.50.417.333 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.50.417.334 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.50.417.334 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.50.417.334 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.50.417.336 I srv get_availabl: prompt cache update took 0.02 ms 0.50.417.353 I slot launch_slot_: id 9 | task 120 | processing task, is_child = 0 0.50.417.355 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 0.50.417.369 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.370 I srv alloc: - prompt is already in the cache, skipping 0.50.417.371 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 0.50.417.385 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.386 I srv alloc: - prompt is already in the cache, skipping 0.50.417.387 I slot get_availabl: id 10 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 0.50.417.388 I srv get_availabl: updating prompt cache 0.50.417.405 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.406 I srv alloc: - prompt is already in the cache, skipping 0.50.417.407 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 0.50.417.407 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.50.417.408 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.50.417.408 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.50.417.410 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.50.417.410 I srv get_availabl: prompt cache update took 0.02 ms 0.50.417.424 I slot launch_slot_: id 10 | task 121 | processing task, is_child = 0 0.50.417.425 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 0.50.417.439 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.440 I srv alloc: - prompt is already in the cache, skipping 0.50.417.442 I slot get_availabl: id 14 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 0.50.417.443 I srv get_availabl: updating prompt cache 0.50.417.457 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.417.458 I srv alloc: - prompt is already in the cache, skipping 0.50.417.458 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 0.50.417.459 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.50.417.459 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.50.417.459 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.50.417.460 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.50.417.460 I srv get_availabl: prompt cache update took 0.02 ms 0.50.417.476 I slot launch_slot_: id 14 | task 122 | processing task, is_child = 0 0.50.420.249 I slot operator(): id 8 | task 119 | Checking checkpoint with [24, 24] against 3... 0.50.420.251 W slot operator(): id 8 | task 119 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.50.420.252 W slot operator(): id 8 | task 119 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 0.50.421.937 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.50.421.968 I slot operator(): id 9 | task 120 | Checking checkpoint with [24, 24] against 28... 0.50.425.485 W slot operator(): id 9 | task 120 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 0.50.425.522 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.50.425.529 I slot operator(): id 10 | task 121 | Checking checkpoint with [24, 24] against 3... 0.50.425.529 W slot operator(): id 10 | task 121 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.50.425.530 W slot operator(): id 10 | task 121 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 0.50.426.968 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.50.426.980 I slot operator(): id 14 | task 122 | Checking checkpoint with [24, 24] against 3... 0.50.426.981 W slot operator(): id 14 | task 122 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.50.426.982 W slot operator(): id 14 | task 122 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 0.50.428.502 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.50.553.473 I slot print_timing: id 3 | task 13 | prompt eval time = 536.00 ms / 29 tokens ( 18.48 ms per token, 54.10 tokens per second) 0.50.553.476 I slot print_timing: id 3 | task 13 | eval time = 8921.76 ms / 128 tokens ( 69.70 ms per token, 14.35 tokens per second) 0.50.553.476 I slot print_timing: id 3 | task 13 | total time = 9457.77 ms / 157 tokens 0.50.553.477 I slot print_timing: id 3 | task 13 | graphs reused = 1 0.50.553.480 I slot print_timing: id 3 | task 13 | draft acceptance = 0.41573 ( 37 accepted / 89 generated), mean acceptance length = 1.42, acceptance rate per position = (0.416) 0.50.553.494 I statistics draft-mtp: #calls(b,g,a) = 27 91 1398, #gen drafts = 1409, #acc drafts = 991, #gen tokens = 1409, #acc tokens = 991, #mean acc len = 1.71, #acc rate/pos = (0.709), dur(b,g,a) = 0.018, 148.123, 1.655 ms 0.50.553.520 I slot release: id 3 | task 13 | stop processing: n_tokens = 156, truncated = 0 0.50.562.628 I srv operator(): Chat format: peg-native 0.50.575.282 I slot create_check: id 8 | task 119 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.50.588.623 I slot create_check: id 10 | task 121 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.50.602.374 I slot create_check: id 14 | task 122 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.50.696.491 I slot get_availabl: id 3 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 0.50.696.493 I srv get_availabl: updating prompt cache 0.50.696.532 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.50.696.535 I srv alloc: - prompt is already in the cache, skipping 0.50.696.535 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 0.50.696.538 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.50.696.539 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.50.696.540 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.50.696.541 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.50.696.542 I srv get_availabl: prompt cache update took 0.05 ms 0.50.696.612 I slot launch_slot_: id 3 | task 125 | processing task, is_child = 0 0.50.698.565 I slot operator(): id 3 | task 125 | Checking checkpoint with [24, 24] against 28... 0.50.702.469 W slot operator(): id 3 | task 125 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 0.50.702.495 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.52.506.359 I slot print_timing: id 4 | task 79 | n_decoded = 101, tg = 19.65 t/s, tg_3s = 19.65 t/s 0.52.510.678 I slot print_timing: id 15 | task 82 | n_decoded = 101, tg = 19.64 t/s, tg_3s = 19.64 t/s 0.52.982.553 I slot print_timing: id 5 | task 78 | n_decoded = 100, tg = 18.33 t/s, tg_3s = 18.33 t/s 0.52.984.972 I slot print_timing: id 12 | task 81 | n_decoded = 100, tg = 18.33 t/s, tg_3s = 18.33 t/s 0.53.172.221 I slot print_timing: id 4 | task 79 | prompt eval time = 157.41 ms / 4 tokens ( 39.35 ms per token, 25.41 tokens per second) 0.53.172.224 I slot print_timing: id 4 | task 79 | eval time = 5805.03 ms / 114 tokens ( 50.92 ms per token, 19.64 tokens per second) 0.53.172.224 I slot print_timing: id 4 | task 79 | total time = 5962.44 ms / 118 tokens 0.53.172.225 I slot print_timing: id 4 | task 79 | graphs reused = 1 0.53.172.228 I slot print_timing: id 4 | task 79 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.53.172.245 I statistics draft-mtp: #calls(b,g,a) = 32 118 1825, #gen drafts = 1836, #acc drafts = 1282, #gen tokens = 1836, #acc tokens = 1282, #mean acc len = 1.70, #acc rate/pos = (0.702), dur(b,g,a) = 0.020, 188.952, 2.144 ms 0.53.172.273 I slot release: id 4 | task 79 | stop processing: n_tokens = 139, truncated = 0 0.53.176.411 I slot print_timing: id 15 | task 82 | prompt eval time = 148.38 ms / 4 tokens ( 37.09 ms per token, 26.96 tokens per second) 0.53.176.414 I slot print_timing: id 15 | task 82 | eval time = 5808.78 ms / 114 tokens ( 50.95 ms per token, 19.63 tokens per second) 0.53.176.415 I slot print_timing: id 15 | task 82 | total time = 5957.16 ms / 118 tokens 0.53.176.415 I slot print_timing: id 15 | task 82 | graphs reused = 1 0.53.176.418 I slot print_timing: id 15 | task 82 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.53.176.424 I statistics draft-mtp: #calls(b,g,a) = 32 118 1836, #gen drafts = 1836, #acc drafts = 1289, #gen tokens = 1836, #acc tokens = 1289, #mean acc len = 1.70, #acc rate/pos = (0.702), dur(b,g,a) = 0.020, 188.952, 2.156 ms 0.53.176.446 I slot release: id 15 | task 82 | stop processing: n_tokens = 139, truncated = 0 0.53.181.194 I srv operator(): Chat format: peg-native 0.53.184.452 I srv operator(): Chat format: peg-native 0.53.261.281 I slot get_availabl: id 4 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.022 0.53.261.283 I srv get_availabl: updating prompt cache 0.53.261.323 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.53.261.325 I srv alloc: - prompt is already in the cache, skipping 0.53.261.326 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.107 0.53.261.329 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.53.261.330 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.53.261.331 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.53.261.331 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.53.261.331 I srv get_availabl: prompt cache update took 0.05 ms 0.53.261.397 I slot launch_slot_: id 4 | task 153 | processing task, is_child = 0 0.53.261.403 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 0.53.261.419 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.53.261.420 I srv alloc: - prompt is already in the cache, skipping 0.53.261.422 I slot get_availabl: id 15 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 0.53.261.422 I srv get_availabl: updating prompt cache 0.53.261.435 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.53.261.436 I srv alloc: - prompt is already in the cache, skipping 0.53.261.437 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 0.53.261.437 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.53.261.438 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.53.261.438 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.53.261.438 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.53.261.438 I srv get_availabl: prompt cache update took 0.02 ms 0.53.261.455 I slot launch_slot_: id 15 | task 154 | processing task, is_child = 0 0.53.263.561 I slot operator(): id 4 | task 153 | Checking checkpoint with [20, 20] against 3... 0.53.263.564 W slot operator(): id 4 | task 153 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.53.263.566 W slot operator(): id 4 | task 153 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 0.53.265.129 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.53.265.157 I slot operator(): id 15 | task 154 | Checking checkpoint with [20, 20] against 24... 0.53.269.017 W slot operator(): id 15 | task 154 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 0.53.269.048 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.53.392.531 I slot create_check: id 4 | task 153 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.54.057.136 I slot print_timing: id 6 | task 100 | n_decoded = 101, tg = 20.25 t/s, tg_3s = 20.25 t/s 0.54.245.078 I slot print_timing: id 1 | task 77 | n_decoded = 101, tg = 15.03 t/s, tg_3s = 15.03 t/s 0.54.247.221 I slot print_timing: id 7 | task 80 | n_decoded = 101, tg = 15.03 t/s, tg_3s = 15.03 t/s 0.54.339.457 I slot print_timing: id 0 | task 98 | n_decoded = 100, tg = 18.48 t/s, tg_3s = 18.48 t/s 0.54.343.394 I slot print_timing: id 11 | task 101 | n_decoded = 100, tg = 18.47 t/s, tg_3s = 18.47 t/s 0.54.717.041 I slot print_timing: id 5 | task 78 | prompt eval time = 314.03 ms / 28 tokens ( 11.22 ms per token, 89.16 tokens per second) 0.54.717.046 I slot print_timing: id 5 | task 78 | eval time = 7189.54 ms / 128 tokens ( 56.17 ms per token, 17.80 tokens per second) 0.54.717.046 I slot print_timing: id 5 | task 78 | total time = 7503.58 ms / 156 tokens 0.54.717.047 I slot print_timing: id 5 | task 78 | graphs reused = 1 0.54.717.050 I slot print_timing: id 5 | task 78 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.54.717.066 I statistics draft-mtp: #calls(b,g,a) = 34 134 2071, #gen drafts = 2085, #acc drafts = 1447, #gen tokens = 2085, #acc tokens = 1447, #mean acc len = 1.70, #acc rate/pos = (0.699), dur(b,g,a) = 0.022, 214.998, 2.425 ms 0.54.717.093 I slot release: id 5 | task 78 | stop processing: n_tokens = 155, truncated = 0 0.54.717.307 I slot print_timing: id 12 | task 81 | prompt eval time = 311.70 ms / 28 tokens ( 11.13 ms per token, 89.83 tokens per second) 0.54.717.309 I slot print_timing: id 12 | task 81 | eval time = 7188.53 ms / 128 tokens ( 56.16 ms per token, 17.81 tokens per second) 0.54.717.309 I slot print_timing: id 12 | task 81 | total time = 7500.24 ms / 156 tokens 0.54.717.309 I slot print_timing: id 12 | task 81 | graphs reused = 1 0.54.717.310 I slot print_timing: id 12 | task 81 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.54.717.312 I statistics draft-mtp: #calls(b,g,a) = 34 134 2071, #gen drafts = 2085, #acc drafts = 1447, #gen tokens = 2085, #acc tokens = 1447, #mean acc len = 1.70, #acc rate/pos = (0.699), dur(b,g,a) = 0.022, 214.998, 2.425 ms 0.54.717.333 I slot release: id 12 | task 81 | stop processing: n_tokens = 155, truncated = 0 0.54.719.907 I slot print_timing: id 6 | task 100 | prompt eval time = 279.56 ms / 25 tokens ( 11.18 ms per token, 89.43 tokens per second) 0.54.719.909 I slot print_timing: id 6 | task 100 | eval time = 5650.37 ms / 114 tokens ( 49.56 ms per token, 20.18 tokens per second) 0.54.719.910 I slot print_timing: id 6 | task 100 | total time = 5929.93 ms / 139 tokens 0.54.719.910 I slot print_timing: id 6 | task 100 | graphs reused = 1 0.54.719.912 I slot print_timing: id 6 | task 100 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.54.719.918 I statistics draft-mtp: #calls(b,g,a) = 34 134 2077, #gen drafts = 2085, #acc drafts = 1452, #gen tokens = 2085, #acc tokens = 1452, #mean acc len = 1.70, #acc rate/pos = (0.699), dur(b,g,a) = 0.022, 214.998, 2.433 ms 0.54.719.938 I slot release: id 6 | task 100 | stop processing: n_tokens = 139, truncated = 0 0.54.727.195 I srv operator(): Chat format: peg-native 0.54.727.847 I srv operator(): Chat format: peg-native 0.54.729.172 I srv operator(): Chat format: peg-native 0.54.802.560 I slot get_availabl: id 5 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 0.54.802.562 I srv get_availabl: updating prompt cache 0.54.802.607 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.54.802.610 I srv alloc: - prompt is already in the cache, skipping 0.54.802.611 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 0.54.802.614 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.54.802.614 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.54.802.615 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.54.802.615 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.54.802.616 I srv get_availabl: prompt cache update took 0.05 ms 0.54.802.682 I slot launch_slot_: id 5 | task 171 | processing task, is_child = 0 0.54.802.684 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.54.802.699 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.54.802.700 I srv alloc: - prompt is already in the cache, skipping 0.54.802.701 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 0.54.802.717 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.54.802.718 I srv alloc: - prompt is already in the cache, skipping 0.54.802.720 I slot get_availabl: id 12 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 0.54.802.720 I srv get_availabl: updating prompt cache 0.54.802.734 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.54.802.735 I srv alloc: - prompt is already in the cache, skipping 0.54.802.735 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 0.54.802.736 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.54.802.737 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.54.802.737 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.54.802.737 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.54.802.737 I srv get_availabl: prompt cache update took 0.02 ms 0.54.802.755 I slot launch_slot_: id 12 | task 172 | processing task, is_child = 0 0.54.802.756 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 0.54.802.769 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.54.802.770 I srv alloc: - prompt is already in the cache, skipping 0.54.802.772 I slot get_availabl: id 6 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 0.54.802.772 I srv get_availabl: updating prompt cache 0.54.802.785 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.54.802.786 I srv alloc: - prompt is already in the cache, skipping 0.54.802.786 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 0.54.802.786 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.54.802.787 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.54.802.787 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.54.802.787 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.54.802.787 I srv get_availabl: prompt cache update took 0.02 ms 0.54.802.802 I slot launch_slot_: id 6 | task 173 | processing task, is_child = 0 0.54.805.264 I slot operator(): id 5 | task 171 | Checking checkpoint with [23, 23] against 3... 0.54.805.267 W slot operator(): id 5 | task 171 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.54.805.270 W slot operator(): id 5 | task 171 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 0.54.806.839 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.54.806.868 I slot operator(): id 6 | task 173 | Checking checkpoint with [20, 20] against 24... 0.54.810.863 W slot operator(): id 6 | task 173 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 0.54.810.897 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.54.810.910 I slot operator(): id 12 | task 172 | Checking checkpoint with [23, 23] against 27... 0.54.814.507 W slot operator(): id 12 | task 172 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 0.54.814.541 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.54.941.138 I slot create_check: id 5 | task 171 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.55.604.643 I slot print_timing: id 8 | task 119 | n_decoded = 101, tg = 20.56 t/s, tg_3s = 20.56 t/s 0.55.606.853 I slot print_timing: id 14 | task 122 | n_decoded = 101, tg = 20.55 t/s, tg_3s = 20.55 t/s 0.55.697.300 I slot print_timing: id 2 | task 99 | n_decoded = 101, tg = 15.24 t/s, tg_3s = 15.24 t/s 0.55.796.379 I slot print_timing: id 13 | task 102 | n_decoded = 101, tg = 15.02 t/s, tg_3s = 15.02 t/s 0.55.985.135 I slot print_timing: id 10 | task 121 | n_decoded = 100, tg = 18.89 t/s, tg_3s = 18.89 t/s 0.56.075.075 I slot print_timing: id 0 | task 98 | prompt eval time = 142.85 ms / 4 tokens ( 35.71 ms per token, 28.00 tokens per second) 0.56.075.078 I slot print_timing: id 0 | task 98 | eval time = 7147.56 ms / 128 tokens ( 55.84 ms per token, 17.91 tokens per second) 0.56.075.079 I slot print_timing: id 0 | task 98 | total time = 7290.41 ms / 132 tokens 0.56.075.079 I slot print_timing: id 0 | task 98 | graphs reused = 1 0.56.075.082 I slot print_timing: id 0 | task 98 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.56.075.098 I statistics draft-mtp: #calls(b,g,a) = 37 148 2286, #gen drafts = 2300, #acc drafts = 1591, #gen tokens = 2300, #acc tokens = 1591, #mean acc len = 1.70, #acc rate/pos = (0.696), dur(b,g,a) = 0.024, 238.355, 2.665 ms 0.56.075.123 I slot release: id 0 | task 98 | stop processing: n_tokens = 155, truncated = 0 0.56.075.329 I slot print_timing: id 11 | task 101 | prompt eval time = 136.30 ms / 4 tokens ( 34.08 ms per token, 29.35 tokens per second) 0.56.075.331 I slot print_timing: id 11 | task 101 | eval time = 7147.55 ms / 128 tokens ( 55.84 ms per token, 17.91 tokens per second) 0.56.075.332 I slot print_timing: id 11 | task 101 | total time = 7283.85 ms / 132 tokens 0.56.075.332 I slot print_timing: id 11 | task 101 | graphs reused = 1 0.56.075.333 I slot print_timing: id 11 | task 101 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.56.075.335 I statistics draft-mtp: #calls(b,g,a) = 37 148 2286, #gen drafts = 2300, #acc drafts = 1591, #gen tokens = 2300, #acc tokens = 1591, #mean acc len = 1.70, #acc rate/pos = (0.696), dur(b,g,a) = 0.024, 238.355, 2.665 ms 0.56.075.355 I slot release: id 11 | task 101 | stop processing: n_tokens = 155, truncated = 0 0.56.085.059 I srv operator(): Chat format: peg-native 0.56.085.119 I srv operator(): Chat format: peg-native 0.56.157.990 I slot print_timing: id 1 | task 77 | prompt eval time = 318.70 ms / 29 tokens ( 10.99 ms per token, 90.99 tokens per second) 0.56.157.993 I slot print_timing: id 1 | task 77 | eval time = 8631.12 ms / 128 tokens ( 67.43 ms per token, 14.83 tokens per second) 0.56.157.993 I slot print_timing: id 1 | task 77 | total time = 8949.82 ms / 157 tokens 0.56.157.994 I slot print_timing: id 1 | task 77 | graphs reused = 1 0.56.157.996 I slot print_timing: id 1 | task 77 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 0.56.158.007 I statistics draft-mtp: #calls(b,g,a) = 37 149 2300, #gen drafts = 2312, #acc drafts = 1599, #gen tokens = 2312, #acc tokens = 1599, #mean acc len = 1.70, #acc rate/pos = (0.695), dur(b,g,a) = 0.024, 240.791, 2.685 ms 0.56.158.031 I slot release: id 1 | task 77 | stop processing: n_tokens = 156, truncated = 0 0.56.158.231 I slot print_timing: id 7 | task 80 | prompt eval time = 313.03 ms / 29 tokens ( 10.79 ms per token, 92.64 tokens per second) 0.56.158.233 I slot print_timing: id 7 | task 80 | eval time = 8630.15 ms / 128 tokens ( 67.42 ms per token, 14.83 tokens per second) 0.56.158.233 I slot print_timing: id 7 | task 80 | total time = 8943.18 ms / 157 tokens 0.56.158.233 I slot print_timing: id 7 | task 80 | graphs reused = 1 0.56.158.234 I slot print_timing: id 7 | task 80 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 0.56.158.236 I statistics draft-mtp: #calls(b,g,a) = 37 149 2300, #gen drafts = 2312, #acc drafts = 1599, #gen tokens = 2312, #acc tokens = 1599, #mean acc len = 1.70, #acc rate/pos = (0.695), dur(b,g,a) = 0.024, 240.791, 2.685 ms 0.56.158.254 I slot release: id 7 | task 80 | stop processing: n_tokens = 156, truncated = 0 0.56.162.925 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 0.56.162.926 I srv get_availabl: updating prompt cache 0.56.162.969 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.56.162.972 I srv alloc: - prompt is already in the cache, skipping 0.56.162.973 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 0.56.162.975 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.56.162.976 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.56.162.977 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.56.162.977 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.56.162.977 I srv get_availabl: prompt cache update took 0.05 ms 0.56.163.041 I slot launch_slot_: id 0 | task 188 | processing task, is_child = 0 0.56.163.043 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 0.56.163.057 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.56.163.058 I srv alloc: - prompt is already in the cache, skipping 0.56.163.058 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.56.163.074 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.56.163.075 I srv alloc: - prompt is already in the cache, skipping 0.56.163.075 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.56.163.091 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.56.163.092 I srv alloc: - prompt is already in the cache, skipping 0.56.163.094 I slot get_availabl: id 1 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 0.56.163.094 I srv get_availabl: updating prompt cache 0.56.163.107 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.56.163.109 I srv alloc: - prompt is already in the cache, skipping 0.56.163.109 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 0.56.163.110 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.56.163.110 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.56.163.110 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.56.163.111 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.56.163.111 I srv get_availabl: prompt cache update took 0.02 ms 0.56.163.129 I slot launch_slot_: id 1 | task 189 | processing task, is_child = 0 0.56.163.131 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.56.163.144 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.56.163.145 I srv alloc: - prompt is already in the cache, skipping 0.56.163.146 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.56.163.159 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.56.163.160 I srv alloc: - prompt is already in the cache, skipping 0.56.165.634 I slot operator(): id 0 | task 188 | Checking checkpoint with [23, 23] against 27... 0.56.167.195 I srv operator(): Chat format: peg-native 0.56.167.559 I srv operator(): Chat format: peg-native 0.56.169.534 W slot operator(): id 0 | task 188 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 0.56.169.563 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.56.169.580 I slot operator(): id 1 | task 189 | Checking checkpoint with [24, 24] against 28... 0.56.173.024 W slot operator(): id 1 | task 189 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 0.56.173.051 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.56.256.660 I slot print_timing: id 8 | task 119 | prompt eval time = 271.12 ms / 25 tokens ( 10.84 ms per token, 92.21 tokens per second) 0.56.256.663 I slot print_timing: id 8 | task 119 | eval time = 5565.28 ms / 114 tokens ( 48.82 ms per token, 20.48 tokens per second) 0.56.256.663 I slot print_timing: id 8 | task 119 | total time = 5836.39 ms / 139 tokens 0.56.256.664 I slot print_timing: id 8 | task 119 | graphs reused = 1 0.56.256.666 I slot print_timing: id 8 | task 119 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.56.256.678 I statistics draft-mtp: #calls(b,g,a) = 39 150 2318, #gen drafts = 2324, #acc drafts = 1610, #gen tokens = 2324, #acc tokens = 1610, #mean acc len = 1.69, #acc rate/pos = (0.695), dur(b,g,a) = 0.025, 243.238, 2.709 ms 0.56.256.704 I slot release: id 8 | task 119 | stop processing: n_tokens = 139, truncated = 0 0.56.258.725 I slot print_timing: id 14 | task 122 | prompt eval time = 264.92 ms / 25 tokens ( 10.60 ms per token, 94.37 tokens per second) 0.56.258.727 I slot print_timing: id 14 | task 122 | eval time = 5566.81 ms / 114 tokens ( 48.83 ms per token, 20.48 tokens per second) 0.56.258.728 I slot print_timing: id 14 | task 122 | total time = 5831.73 ms / 139 tokens 0.56.258.728 I slot print_timing: id 14 | task 122 | graphs reused = 1 0.56.258.731 I slot print_timing: id 14 | task 122 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.56.258.737 I statistics draft-mtp: #calls(b,g,a) = 39 150 2323, #gen drafts = 2324, #acc drafts = 1614, #gen tokens = 2324, #acc tokens = 1614, #mean acc len = 1.69, #acc rate/pos = (0.695), dur(b,g,a) = 0.025, 243.238, 2.716 ms 0.56.258.759 I slot release: id 14 | task 122 | stop processing: n_tokens = 139, truncated = 0 0.56.259.210 I slot get_availabl: id 8 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 0.56.259.211 I srv get_availabl: updating prompt cache 0.56.259.251 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.56.259.253 I srv alloc: - prompt is already in the cache, skipping 0.56.259.254 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 0.56.259.257 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.56.259.258 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.56.259.258 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.56.259.258 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.56.259.260 I srv get_availabl: prompt cache update took 0.05 ms 0.56.259.323 I slot launch_slot_: id 8 | task 191 | processing task, is_child = 0 0.56.259.324 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 0.56.259.340 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.56.259.342 I srv alloc: - prompt is already in the cache, skipping 0.56.259.342 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.56.259.357 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.56.259.359 I srv alloc: - prompt is already in the cache, skipping 0.56.259.359 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 0.56.259.374 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.56.259.376 I srv alloc: - prompt is already in the cache, skipping 0.56.259.378 I slot get_availabl: id 7 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 0.56.259.378 I srv get_availabl: updating prompt cache 0.56.259.391 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.56.259.392 I srv alloc: - prompt is already in the cache, skipping 0.56.259.393 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 0.56.259.393 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.56.259.394 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.56.259.395 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.56.259.395 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.56.259.395 I srv get_availabl: prompt cache update took 0.02 ms 0.56.259.419 I slot launch_slot_: id 7 | task 192 | processing task, is_child = 0 0.56.259.421 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 0.56.259.434 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.56.259.435 I srv alloc: - prompt is already in the cache, skipping 0.56.259.436 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 0.56.259.449 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.56.259.450 I srv alloc: - prompt is already in the cache, skipping 0.56.262.205 I slot operator(): id 7 | task 192 | Checking checkpoint with [24, 24] against 28... 0.56.265.484 I srv operator(): Chat format: peg-native 0.56.266.048 W slot operator(): id 7 | task 192 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 0.56.266.077 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.56.266.095 I slot operator(): id 8 | task 191 | Checking checkpoint with [20, 20] against 24... 0.56.266.910 I srv operator(): Chat format: peg-native 0.56.269.599 W slot operator(): id 8 | task 191 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 0.56.269.633 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.56.355.831 I slot get_availabl: id 11 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 0.56.355.834 I srv get_availabl: updating prompt cache 0.56.355.874 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.56.355.877 I srv alloc: - prompt is already in the cache, skipping 0.56.355.878 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 0.56.355.881 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.56.355.881 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.56.355.882 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.56.355.883 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.56.355.885 I srv get_availabl: prompt cache update took 0.05 ms 0.56.355.951 I slot launch_slot_: id 11 | task 194 | processing task, is_child = 0 0.56.355.953 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 0.56.355.969 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.56.355.970 I srv alloc: - prompt is already in the cache, skipping 0.56.355.972 I slot get_availabl: id 14 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 0.56.355.972 I srv get_availabl: updating prompt cache 0.56.355.985 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.56.355.987 I srv alloc: - prompt is already in the cache, skipping 0.56.355.987 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 0.56.355.988 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.56.355.988 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.56.355.989 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.56.355.989 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.56.355.990 I srv get_availabl: prompt cache update took 0.02 ms 0.56.356.011 I slot launch_slot_: id 14 | task 195 | processing task, is_child = 0 0.56.358.341 I slot operator(): id 11 | task 194 | Checking checkpoint with [23, 23] against 27... 0.56.362.249 W slot operator(): id 11 | task 194 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 0.56.362.276 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.56.362.297 I slot operator(): id 14 | task 195 | Checking checkpoint with [20, 20] against 24... 0.56.365.840 W slot operator(): id 14 | task 195 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 0.56.365.872 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.57.318.497 I slot print_timing: id 9 | task 120 | n_decoded = 101, tg = 14.93 t/s, tg_3s = 14.93 t/s 0.57.507.080 I slot print_timing: id 3 | task 125 | n_decoded = 101, tg = 15.04 t/s, tg_3s = 15.04 t/s 0.57.602.447 I slot print_timing: id 2 | task 99 | prompt eval time = 281.04 ms / 29 tokens ( 9.69 ms per token, 103.19 tokens per second) 0.57.602.450 I slot print_timing: id 2 | task 99 | eval time = 8533.16 ms / 128 tokens ( 66.67 ms per token, 15.00 tokens per second) 0.57.602.450 I slot print_timing: id 2 | task 99 | total time = 8814.20 ms / 157 tokens 0.57.602.451 I slot print_timing: id 2 | task 99 | graphs reused = 1 0.57.602.453 I slot print_timing: id 2 | task 99 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 0.57.602.469 I statistics draft-mtp: #calls(b,g,a) = 43 164 2526, #gen drafts = 2541, #acc drafts = 1752, #gen tokens = 2541, #acc tokens = 1752, #mean acc len = 1.69, #acc rate/pos = (0.694), dur(b,g,a) = 0.027, 267.677, 2.953 ms 0.57.602.498 I slot release: id 2 | task 99 | stop processing: n_tokens = 156, truncated = 0 0.57.612.213 I srv operator(): Chat format: peg-native 0.57.691.898 I slot print_timing: id 10 | task 121 | prompt eval time = 266.12 ms / 28 tokens ( 9.50 ms per token, 105.22 tokens per second) 0.57.691.901 I slot print_timing: id 10 | task 121 | eval time = 7000.22 ms / 128 tokens ( 54.69 ms per token, 18.29 tokens per second) 0.57.691.901 I slot print_timing: id 10 | task 121 | total time = 7266.34 ms / 156 tokens 0.57.691.902 I slot print_timing: id 10 | task 121 | graphs reused = 1 0.57.691.905 I slot print_timing: id 10 | task 121 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 0.57.691.919 I statistics draft-mtp: #calls(b,g,a) = 43 165 2541, #gen drafts = 2554, #acc drafts = 1762, #gen tokens = 2554, #acc tokens = 1762, #mean acc len = 1.69, #acc rate/pos = (0.693), dur(b,g,a) = 0.027, 270.521, 2.972 ms 0.57.691.944 I slot release: id 10 | task 121 | stop processing: n_tokens = 155, truncated = 0 0.57.692.150 I slot print_timing: id 13 | task 102 | prompt eval time = 274.94 ms / 29 tokens ( 9.48 ms per token, 105.48 tokens per second) 0.57.692.152 I slot print_timing: id 13 | task 102 | eval time = 8622.17 ms / 128 tokens ( 67.36 ms per token, 14.85 tokens per second) 0.57.692.153 I slot print_timing: id 13 | task 102 | total time = 8897.10 ms / 157 tokens 0.57.692.153 I slot print_timing: id 13 | task 102 | graphs reused = 1 0.57.692.154 I slot print_timing: id 13 | task 102 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 0.57.692.156 I statistics draft-mtp: #calls(b,g,a) = 43 165 2541, #gen drafts = 2554, #acc drafts = 1762, #gen tokens = 2554, #acc tokens = 1762, #mean acc len = 1.69, #acc rate/pos = (0.693), dur(b,g,a) = 0.027, 270.521, 2.972 ms 0.57.692.170 I slot release: id 13 | task 102 | stop processing: n_tokens = 156, truncated = 0 0.57.697.591 I slot get_availabl: id 2 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 0.57.697.593 I srv get_availabl: updating prompt cache 0.57.697.635 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.57.697.638 I srv alloc: - prompt is already in the cache, skipping 0.57.697.638 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 0.57.697.641 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.57.697.642 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.57.697.642 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.57.697.643 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.57.697.643 I srv get_availabl: prompt cache update took 0.05 ms 0.57.697.705 I slot launch_slot_: id 2 | task 210 | processing task, is_child = 0 0.57.697.707 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 0.57.697.723 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.57.697.724 I srv alloc: - prompt is already in the cache, skipping 0.57.697.725 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 0.57.697.740 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.57.697.742 I srv alloc: - prompt is already in the cache, skipping 0.57.700.574 I slot operator(): id 2 | task 210 | Checking checkpoint with [24, 24] against 28... 0.57.701.997 I srv operator(): Chat format: peg-native 0.57.702.110 I srv operator(): Chat format: peg-native 0.57.704.404 W slot operator(): id 2 | task 210 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 0.57.704.453 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.57.789.675 I slot get_availabl: id 10 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 0.57.789.677 I srv get_availabl: updating prompt cache 0.57.789.717 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 0.57.789.719 I srv alloc: - prompt is already in the cache, skipping 0.57.789.720 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 0.57.789.722 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.57.789.723 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.57.789.724 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.57.789.725 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.57.789.726 I srv get_availabl: prompt cache update took 0.05 ms 0.57.789.791 I slot launch_slot_: id 10 | task 212 | processing task, is_child = 0 0.57.789.792 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 0.57.789.808 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.57.789.809 I srv alloc: - prompt is already in the cache, skipping 0.57.789.811 I slot get_availabl: id 13 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 0.57.789.811 I srv get_availabl: updating prompt cache 0.57.789.825 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.57.789.826 I srv alloc: - prompt is already in the cache, skipping 0.57.789.827 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 0.57.789.828 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.57.789.828 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.57.789.828 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.57.789.829 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.57.789.829 I srv get_availabl: prompt cache update took 0.02 ms 0.57.789.847 I slot launch_slot_: id 13 | task 213 | processing task, is_child = 0 0.57.792.139 I slot operator(): id 10 | task 212 | Checking checkpoint with [23, 23] against 27... 0.57.795.960 W slot operator(): id 10 | task 212 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 0.57.795.994 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.57.796.007 I slot operator(): id 13 | task 213 | Checking checkpoint with [24, 24] against 3... 0.57.796.009 W slot operator(): id 13 | task 213 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.57.796.010 W slot operator(): id 13 | task 213 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 0.57.797.941 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.57.921.420 I slot create_check: id 13 | task 213 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.58.308.846 I slot print_timing: id 15 | task 154 | n_decoded = 101, tg = 20.46 t/s, tg_3s = 20.46 t/s 0.58.785.222 I slot print_timing: id 4 | task 153 | n_decoded = 100, tg = 18.86 t/s, tg_3s = 18.86 t/s 0.58.980.496 I slot print_timing: id 15 | task 154 | prompt eval time = 106.42 ms / 4 tokens ( 26.60 ms per token, 37.59 tokens per second) 0.58.980.499 I slot print_timing: id 15 | task 154 | eval time = 5608.90 ms / 114 tokens ( 49.20 ms per token, 20.32 tokens per second) 0.58.980.499 I slot print_timing: id 15 | task 154 | total time = 5715.32 ms / 118 tokens 0.58.980.500 I slot print_timing: id 15 | task 154 | graphs reused = 1 0.58.980.503 I slot print_timing: id 15 | task 154 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 0.58.980.515 I statistics draft-mtp: #calls(b,g,a) = 46 178 2756, #gen drafts = 2756, #acc drafts = 1916, #gen tokens = 2756, #acc tokens = 1916, #mean acc len = 1.70, #acc rate/pos = (0.695), dur(b,g,a) = 0.029, 292.014, 3.230 ms 0.58.980.541 I slot release: id 15 | task 154 | stop processing: n_tokens = 139, truncated = 0 0.58.989.594 I srv operator(): Chat format: peg-native 0.59.072.262 I slot get_availabl: id 15 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 0.59.072.265 I srv get_availabl: updating prompt cache 0.59.072.305 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 0.59.072.307 I srv alloc: - prompt is already in the cache, skipping 0.59.072.308 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 0.59.072.311 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.59.072.312 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.59.072.313 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.59.072.314 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.59.072.314 I srv get_availabl: prompt cache update took 0.05 ms 0.59.072.383 I slot launch_slot_: id 15 | task 227 | processing task, is_child = 0 0.59.073.824 I slot operator(): id 15 | task 227 | Checking checkpoint with [20, 20] against 3... 0.59.073.826 W slot operator(): id 15 | task 227 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.59.073.828 W slot operator(): id 15 | task 227 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 0.59.075.413 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.59.197.951 I slot create_check: id 15 | task 227 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 0.59.286.902 I slot print_timing: id 9 | task 120 | prompt eval time = 131.81 ms / 4 tokens ( 32.95 ms per token, 30.35 tokens per second) 0.59.286.905 I slot print_timing: id 9 | task 120 | eval time = 8733.09 ms / 128 tokens ( 68.23 ms per token, 14.66 tokens per second) 0.59.286.906 I slot print_timing: id 9 | task 120 | total time = 8864.90 ms / 132 tokens 0.59.286.906 I slot print_timing: id 9 | task 120 | graphs reused = 1 0.59.286.908 I slot print_timing: id 9 | task 120 | draft acceptance = 0.41573 ( 37 accepted / 89 generated), mean acceptance length = 1.42, acceptance rate per position = (0.416) 0.59.286.922 I statistics draft-mtp: #calls(b,g,a) = 46 181 2786, #gen drafts = 2800, #acc drafts = 1941, #gen tokens = 2800, #acc tokens = 1941, #mean acc len = 1.70, #acc rate/pos = (0.697), dur(b,g,a) = 0.029, 298.550, 3.266 ms 0.59.286.950 I slot release: id 9 | task 120 | stop processing: n_tokens = 156, truncated = 0 0.59.295.997 I srv operator(): Chat format: peg-native 0.59.383.568 I slot get_availabl: id 9 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 0.59.383.571 I srv get_availabl: updating prompt cache 0.59.383.611 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.59.383.613 I srv alloc: - prompt is already in the cache, skipping 0.59.383.614 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 0.59.383.617 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.59.383.617 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.59.383.619 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.59.383.620 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.59.383.620 I srv get_availabl: prompt cache update took 0.05 ms 0.59.383.685 I slot launch_slot_: id 9 | task 231 | processing task, is_child = 0 0.59.386.253 I slot operator(): id 9 | task 231 | Checking checkpoint with [24, 24] against 3... 0.59.386.256 W slot operator(): id 9 | task 231 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.59.386.258 W slot operator(): id 9 | task 231 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 0.59.387.750 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.59.486.698 I slot print_timing: id 3 | task 125 | prompt eval time = 93.59 ms / 4 tokens ( 23.40 ms per token, 42.74 tokens per second) 0.59.486.701 I slot print_timing: id 3 | task 125 | eval time = 8694.51 ms / 128 tokens ( 67.93 ms per token, 14.72 tokens per second) 0.59.486.701 I slot print_timing: id 3 | task 125 | total time = 8788.11 ms / 132 tokens 0.59.486.702 I slot print_timing: id 3 | task 125 | graphs reused = 1 0.59.486.705 I slot print_timing: id 3 | task 125 | draft acceptance = 0.41573 ( 37 accepted / 89 generated), mean acceptance length = 1.42, acceptance rate per position = (0.416) 0.59.486.718 I statistics draft-mtp: #calls(b,g,a) = 47 183 2815, #gen drafts = 2829, #acc drafts = 1959, #gen tokens = 2829, #acc tokens = 1959, #mean acc len = 1.70, #acc rate/pos = (0.696), dur(b,g,a) = 0.030, 303.192, 3.301 ms 0.59.486.748 I slot release: id 3 | task 125 | stop processing: n_tokens = 156, truncated = 0 0.59.495.889 I srv operator(): Chat format: peg-native 0.59.508.439 I slot create_check: id 9 | task 231 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 0.59.599.285 I slot get_availabl: id 3 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 0.59.599.287 I srv get_availabl: updating prompt cache 0.59.599.326 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 0.59.599.328 I srv alloc: - prompt is already in the cache, skipping 0.59.599.329 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 0.59.599.332 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 0.59.599.333 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 0.59.599.334 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 0.59.599.334 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 0.59.599.334 I srv get_availabl: prompt cache update took 0.05 ms 0.59.599.392 I slot launch_slot_: id 3 | task 234 | processing task, is_child = 0 0.59.601.414 I slot operator(): id 3 | task 234 | Checking checkpoint with [24, 24] against 3... 0.59.601.416 W slot operator(): id 3 | task 234 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 0.59.601.419 W slot operator(): id 3 | task 234 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 0.59.602.976 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 0.59.723.966 I slot create_check: id 3 | task 234 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 0.59.912.456 I slot print_timing: id 6 | task 173 | n_decoded = 101, tg = 20.23 t/s, tg_3s = 20.23 t/s 1.00.296.101 I slot print_timing: id 12 | task 172 | n_decoded = 100, tg = 18.60 t/s, tg_3s = 18.60 t/s 1.00.576.522 I slot print_timing: id 4 | task 153 | prompt eval time = 218.97 ms / 28 tokens ( 7.82 ms per token, 127.87 tokens per second) 1.00.576.525 I slot print_timing: id 4 | task 153 | eval time = 7093.97 ms / 128 tokens ( 55.42 ms per token, 18.04 tokens per second) 1.00.576.526 I slot print_timing: id 4 | task 153 | total time = 7312.94 ms / 156 tokens 1.00.576.527 I slot print_timing: id 4 | task 153 | graphs reused = 1 1.00.576.529 I slot print_timing: id 4 | task 153 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.00.576.543 I statistics draft-mtp: #calls(b,g,a) = 49 194 2985, #gen drafts = 3000, #acc drafts = 2077, #gen tokens = 3000, #acc tokens = 2077, #mean acc len = 1.70, #acc rate/pos = (0.696), dur(b,g,a) = 0.032, 321.982, 3.502 ms 1.00.576.570 I slot release: id 4 | task 153 | stop processing: n_tokens = 155, truncated = 0 1.00.579.279 I slot print_timing: id 6 | task 173 | prompt eval time = 113.11 ms / 4 tokens ( 28.28 ms per token, 35.37 tokens per second) 1.00.579.281 I slot print_timing: id 6 | task 173 | eval time = 5659.29 ms / 114 tokens ( 49.64 ms per token, 20.14 tokens per second) 1.00.579.282 I slot print_timing: id 6 | task 173 | total time = 5772.39 ms / 118 tokens 1.00.579.282 I slot print_timing: id 6 | task 173 | graphs reused = 1 1.00.579.284 I slot print_timing: id 6 | task 173 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.00.579.289 I statistics draft-mtp: #calls(b,g,a) = 49 194 2991, #gen drafts = 3000, #acc drafts = 2083, #gen tokens = 3000, #acc tokens = 2083, #mean acc len = 1.70, #acc rate/pos = (0.696), dur(b,g,a) = 0.032, 321.982, 3.509 ms 1.00.579.309 I slot release: id 6 | task 173 | stop processing: n_tokens = 139, truncated = 0 1.00.585.956 I srv operator(): Chat format: peg-native 1.00.588.192 I srv operator(): Chat format: peg-native 1.00.668.181 I slot get_availabl: id 4 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 1.00.668.183 I srv get_availabl: updating prompt cache 1.00.668.227 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.00.668.230 I srv alloc: - prompt is already in the cache, skipping 1.00.668.230 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 1.00.668.233 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.00.668.234 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.00.668.235 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.00.668.236 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.00.668.237 I srv get_availabl: prompt cache update took 0.05 ms 1.00.668.304 I slot launch_slot_: id 4 | task 246 | processing task, is_child = 0 1.00.668.305 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 1.00.668.321 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.00.668.323 I srv alloc: - prompt is already in the cache, skipping 1.00.668.325 I slot get_availabl: id 6 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.022 1.00.668.325 I srv get_availabl: updating prompt cache 1.00.668.339 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.00.668.340 I srv alloc: - prompt is already in the cache, skipping 1.00.668.340 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.107 1.00.668.341 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.00.668.341 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.00.668.342 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.00.668.342 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.00.668.343 I srv get_availabl: prompt cache update took 0.02 ms 1.00.668.359 I slot launch_slot_: id 6 | task 247 | processing task, is_child = 0 1.00.670.859 I slot operator(): id 4 | task 246 | Checking checkpoint with [23, 23] against 3... 1.00.670.862 W slot operator(): id 4 | task 246 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.00.670.863 W slot operator(): id 4 | task 246 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.00.672.511 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.00.672.540 I slot operator(): id 6 | task 247 | Checking checkpoint with [20, 20] against 3... 1.00.672.542 W slot operator(): id 6 | task 247 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.00.672.543 W slot operator(): id 6 | task 247 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.00.674.344 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.00.809.193 I slot create_check: id 4 | task 246 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.00.823.655 I slot create_check: id 6 | task 247 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.01.395.664 I slot print_timing: id 8 | task 191 | n_decoded = 101, tg = 20.02 t/s, tg_3s = 20.02 t/s 1.01.493.118 I slot print_timing: id 14 | task 195 | n_decoded = 101, tg = 20.06 t/s, tg_3s = 20.06 t/s 1.01.677.994 I slot print_timing: id 0 | task 188 | n_decoded = 100, tg = 18.44 t/s, tg_3s = 18.44 t/s 1.01.680.095 I slot print_timing: id 5 | task 171 | n_decoded = 101, tg = 15.19 t/s, tg_3s = 15.19 t/s 1.01.872.932 I slot print_timing: id 11 | task 194 | n_decoded = 100, tg = 18.47 t/s, tg_3s = 18.47 t/s 1.02.059.211 I slot print_timing: id 12 | task 172 | prompt eval time = 109.34 ms / 4 tokens ( 27.34 ms per token, 36.58 tokens per second) 1.02.059.215 I slot print_timing: id 12 | task 172 | eval time = 7138.93 ms / 128 tokens ( 55.77 ms per token, 17.93 tokens per second) 1.02.059.215 I slot print_timing: id 12 | task 172 | total time = 7248.27 ms / 132 tokens 1.02.059.216 I slot print_timing: id 12 | task 172 | graphs reused = 1 1.02.059.219 I slot print_timing: id 12 | task 172 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.02.059.232 I statistics draft-mtp: #calls(b,g,a) = 51 209 3218, #gen drafts = 3233, #acc drafts = 2239, #gen tokens = 3233, #acc tokens = 2239, #mean acc len = 1.70, #acc rate/pos = (0.696), dur(b,g,a) = 0.033, 347.430, 3.769 ms 1.02.059.263 I slot release: id 12 | task 172 | stop processing: n_tokens = 155, truncated = 0 1.02.062.914 I slot print_timing: id 8 | task 191 | prompt eval time = 84.95 ms / 4 tokens ( 21.24 ms per token, 47.09 tokens per second) 1.02.062.916 I slot print_timing: id 8 | task 191 | eval time = 5711.85 ms / 114 tokens ( 50.10 ms per token, 19.96 tokens per second) 1.02.062.917 I slot print_timing: id 8 | task 191 | total time = 5796.80 ms / 118 tokens 1.02.062.917 I slot print_timing: id 8 | task 191 | graphs reused = 1 1.02.062.920 I slot print_timing: id 8 | task 191 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.02.062.927 I statistics draft-mtp: #calls(b,g,a) = 51 209 3227, #gen drafts = 3233, #acc drafts = 2245, #gen tokens = 3233, #acc tokens = 2245, #mean acc len = 1.70, #acc rate/pos = (0.696), dur(b,g,a) = 0.033, 347.430, 3.779 ms 1.02.062.951 I slot release: id 8 | task 191 | stop processing: n_tokens = 139, truncated = 0 1.02.068.726 I srv operator(): Chat format: peg-native 1.02.071.073 I srv operator(): Chat format: peg-native 1.02.150.806 I slot print_timing: id 14 | task 195 | prompt eval time = 95.76 ms / 4 tokens ( 23.94 ms per token, 41.77 tokens per second) 1.02.150.809 I slot print_timing: id 14 | task 195 | eval time = 5692.72 ms / 114 tokens ( 49.94 ms per token, 20.03 tokens per second) 1.02.150.809 I slot print_timing: id 14 | task 195 | total time = 5788.48 ms / 118 tokens 1.02.150.810 I slot print_timing: id 14 | task 195 | graphs reused = 1 1.02.150.812 I slot print_timing: id 14 | task 195 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.02.150.824 I statistics draft-mtp: #calls(b,g,a) = 51 210 3246, #gen drafts = 3247, #acc drafts = 2258, #gen tokens = 3247, #acc tokens = 2258, #mean acc len = 1.70, #acc rate/pos = (0.696), dur(b,g,a) = 0.033, 350.009, 3.799 ms 1.02.150.850 I slot release: id 14 | task 195 | stop processing: n_tokens = 139, truncated = 0 1.02.151.307 I slot get_availabl: id 8 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.02.151.309 I srv get_availabl: updating prompt cache 1.02.151.350 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.02.151.352 I srv alloc: - prompt is already in the cache, skipping 1.02.151.353 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.02.151.356 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.02.151.357 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.02.151.357 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.02.151.357 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.02.151.358 I srv get_availabl: prompt cache update took 0.05 ms 1.02.151.429 I slot launch_slot_: id 8 | task 263 | processing task, is_child = 0 1.02.151.430 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 1.02.151.446 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.02.151.448 I srv alloc: - prompt is already in the cache, skipping 1.02.151.448 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 1.02.151.463 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.02.151.464 I srv alloc: - prompt is already in the cache, skipping 1.02.151.466 I slot get_availabl: id 12 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 1.02.151.466 I srv get_availabl: updating prompt cache 1.02.151.480 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.02.151.481 I srv alloc: - prompt is already in the cache, skipping 1.02.151.482 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 1.02.151.482 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.02.151.483 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.02.151.483 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.02.151.483 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.02.151.484 I srv get_availabl: prompt cache update took 0.02 ms 1.02.151.502 I slot launch_slot_: id 12 | task 264 | processing task, is_child = 0 1.02.151.503 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 1.02.151.517 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.02.151.518 I srv alloc: - prompt is already in the cache, skipping 1.02.154.418 I slot operator(): id 8 | task 263 | Checking checkpoint with [20, 20] against 24... 1.02.158.275 W slot operator(): id 8 | task 263 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.02.158.310 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.02.158.327 I slot operator(): id 12 | task 264 | Checking checkpoint with [23, 23] against 3... 1.02.158.328 W slot operator(): id 12 | task 264 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.02.158.329 W slot operator(): id 12 | task 264 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.02.159.025 I srv operator(): Chat format: peg-native 1.02.159.965 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.02.263.228 I slot get_availabl: id 14 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.022 1.02.263.230 I srv get_availabl: updating prompt cache 1.02.263.271 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.02.263.273 I srv alloc: - prompt is already in the cache, skipping 1.02.263.274 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.107 1.02.263.277 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.02.263.278 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.02.263.278 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.02.263.280 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.02.263.280 I srv get_availabl: prompt cache update took 0.05 ms 1.02.263.348 I slot launch_slot_: id 14 | task 266 | processing task, is_child = 0 1.02.279.097 I slot create_check: id 12 | task 264 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.02.279.103 I slot operator(): id 14 | task 266 | Checking checkpoint with [20, 20] against 3... 1.02.279.103 W slot operator(): id 14 | task 266 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.02.279.104 W slot operator(): id 14 | task 266 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.02.280.738 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.02.403.971 I slot create_check: id 14 | task 266 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.03.066.892 I slot print_timing: id 1 | task 189 | n_decoded = 101, tg = 14.83 t/s, tg_3s = 14.83 t/s 1.03.071.443 I slot print_timing: id 13 | task 213 | n_decoded = 101, tg = 19.96 t/s, tg_3s = 19.96 t/s 1.03.259.464 I slot print_timing: id 7 | task 192 | n_decoded = 101, tg = 14.62 t/s, tg_3s = 14.62 t/s 1.03.356.054 I slot print_timing: id 10 | task 212 | n_decoded = 100, tg = 18.33 t/s, tg_3s = 18.33 t/s 1.03.446.195 I slot print_timing: id 0 | task 188 | prompt eval time = 88.47 ms / 4 tokens ( 22.12 ms per token, 45.21 tokens per second) 1.03.446.199 I slot print_timing: id 0 | task 188 | eval time = 7192.06 ms / 128 tokens ( 56.19 ms per token, 17.80 tokens per second) 1.03.446.199 I slot print_timing: id 0 | task 188 | total time = 7280.53 ms / 132 tokens 1.03.446.200 I slot print_timing: id 0 | task 188 | graphs reused = 1 1.03.446.203 I slot print_timing: id 0 | task 188 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.03.446.217 I statistics draft-mtp: #calls(b,g,a) = 54 223 3433, #gen drafts = 3448, #acc drafts = 2378, #gen tokens = 3448, #acc tokens = 2378, #mean acc len = 1.69, #acc rate/pos = (0.693), dur(b,g,a) = 0.037, 371.581, 4.032 ms 1.03.446.247 I slot release: id 0 | task 188 | stop processing: n_tokens = 155, truncated = 0 1.03.455.447 I srv operator(): Chat format: peg-native 1.03.541.389 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.03.541.392 I srv get_availabl: updating prompt cache 1.03.541.438 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.03.541.440 I srv alloc: - prompt is already in the cache, skipping 1.03.541.441 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.03.541.444 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.03.541.445 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.03.541.445 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.03.541.446 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.03.541.448 I srv get_availabl: prompt cache update took 0.06 ms 1.03.541.518 I slot launch_slot_: id 0 | task 280 | processing task, is_child = 0 1.03.544.213 I slot operator(): id 0 | task 280 | Checking checkpoint with [23, 23] against 3... 1.03.544.215 W slot operator(): id 0 | task 280 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.03.544.218 W slot operator(): id 0 | task 280 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.03.545.825 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.03.643.163 I slot print_timing: id 5 | task 171 | prompt eval time = 225.75 ms / 29 tokens ( 7.78 ms per token, 128.46 tokens per second) 1.03.643.166 I slot print_timing: id 5 | task 171 | eval time = 8612.12 ms / 128 tokens ( 67.28 ms per token, 14.86 tokens per second) 1.03.643.166 I slot print_timing: id 5 | task 171 | total time = 8837.87 ms / 157 tokens 1.03.643.167 I slot print_timing: id 5 | task 171 | graphs reused = 1 1.03.643.170 I slot print_timing: id 5 | task 171 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.03.643.183 I statistics draft-mtp: #calls(b,g,a) = 54 225 3463, #gen drafts = 3476, #acc drafts = 2397, #gen tokens = 3476, #acc tokens = 2397, #mean acc len = 1.69, #acc rate/pos = (0.692), dur(b,g,a) = 0.037, 375.685, 4.068 ms 1.03.643.212 I slot release: id 5 | task 171 | stop processing: n_tokens = 156, truncated = 0 1.03.643.429 I slot print_timing: id 11 | task 194 | prompt eval time = 99.45 ms / 4 tokens ( 24.86 ms per token, 40.22 tokens per second) 1.03.643.431 I slot print_timing: id 11 | task 194 | eval time = 7185.63 ms / 128 tokens ( 56.14 ms per token, 17.81 tokens per second) 1.03.643.431 I slot print_timing: id 11 | task 194 | total time = 7285.08 ms / 132 tokens 1.03.643.431 I slot print_timing: id 11 | task 194 | graphs reused = 1 1.03.643.432 I slot print_timing: id 11 | task 194 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.03.643.434 I statistics draft-mtp: #calls(b,g,a) = 54 225 3463, #gen drafts = 3476, #acc drafts = 2397, #gen tokens = 3476, #acc tokens = 2397, #mean acc len = 1.69, #acc rate/pos = (0.692), dur(b,g,a) = 0.037, 375.685, 4.068 ms 1.03.643.463 I slot release: id 11 | task 194 | stop processing: n_tokens = 155, truncated = 0 1.03.652.687 I srv operator(): Chat format: peg-native 1.03.652.968 I srv operator(): Chat format: peg-native 1.03.664.645 I slot create_check: id 0 | task 280 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.03.748.317 I slot print_timing: id 13 | task 213 | prompt eval time = 215.19 ms / 25 tokens ( 8.61 ms per token, 116.18 tokens per second) 1.03.748.320 I slot print_timing: id 13 | task 213 | eval time = 5737.10 ms / 114 tokens ( 50.33 ms per token, 19.87 tokens per second) 1.03.748.321 I slot print_timing: id 13 | task 213 | total time = 5952.29 ms / 139 tokens 1.03.748.322 I slot print_timing: id 13 | task 213 | graphs reused = 1 1.03.748.324 I slot print_timing: id 13 | task 213 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.03.748.338 I statistics draft-mtp: #calls(b,g,a) = 55 226 3487, #gen drafts = 3489, #acc drafts = 2412, #gen tokens = 3489, #acc tokens = 2412, #mean acc len = 1.69, #acc rate/pos = (0.692), dur(b,g,a) = 0.038, 378.387, 4.099 ms 1.03.748.365 I slot release: id 13 | task 213 | stop processing: n_tokens = 139, truncated = 0 1.03.749.130 I slot get_availabl: id 11 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.03.749.132 I srv get_availabl: updating prompt cache 1.03.749.174 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.03.749.177 I srv alloc: - prompt is already in the cache, skipping 1.03.749.177 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.03.749.181 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.03.749.182 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.03.749.183 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.03.749.183 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.03.749.183 I srv get_availabl: prompt cache update took 0.05 ms 1.03.749.252 I slot launch_slot_: id 11 | task 283 | processing task, is_child = 0 1.03.749.254 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 1.03.749.271 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.03.749.272 I srv alloc: - prompt is already in the cache, skipping 1.03.749.272 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 1.03.749.287 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.03.749.289 I srv alloc: - prompt is already in the cache, skipping 1.03.749.291 I slot get_availabl: id 5 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.03.749.291 I srv get_availabl: updating prompt cache 1.03.749.305 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.03.749.306 I srv alloc: - prompt is already in the cache, skipping 1.03.749.306 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.03.749.307 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.03.749.307 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.03.749.308 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.03.749.309 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.03.749.309 I srv get_availabl: prompt cache update took 0.02 ms 1.03.749.327 I slot launch_slot_: id 5 | task 284 | processing task, is_child = 0 1.03.749.329 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 1.03.749.342 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.03.749.343 I srv alloc: - prompt is already in the cache, skipping 1.03.752.171 I slot operator(): id 5 | task 284 | Checking checkpoint with [24, 24] against 28... 1.03.756.101 W slot operator(): id 5 | task 284 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.03.756.136 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.03.756.151 I slot operator(): id 11 | task 283 | Checking checkpoint with [23, 23] against 27... 1.03.756.692 I srv operator(): Chat format: peg-native 1.03.759.642 W slot operator(): id 11 | task 283 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.03.759.671 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.03.852.566 I slot get_availabl: id 13 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.03.852.569 I srv get_availabl: updating prompt cache 1.03.852.609 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.03.852.612 I srv alloc: - prompt is already in the cache, skipping 1.03.852.612 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.03.852.615 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.03.852.616 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.03.852.617 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.03.852.618 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.03.852.619 I srv get_availabl: prompt cache update took 0.05 ms 1.03.852.687 I slot launch_slot_: id 13 | task 286 | processing task, is_child = 0 1.03.854.619 I slot operator(): id 13 | task 286 | Checking checkpoint with [20, 20] against 24... 1.03.858.569 W slot operator(): id 13 | task 286 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.03.858.598 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.04.521.199 I slot print_timing: id 2 | task 210 | n_decoded = 101, tg = 14.99 t/s, tg_3s = 14.99 t/s 1.04.807.880 I slot print_timing: id 3 | task 234 | n_decoded = 101, tg = 20.22 t/s, tg_3s = 20.22 t/s 1.04.997.172 I slot print_timing: id 1 | task 189 | prompt eval time = 84.80 ms / 4 tokens ( 21.20 ms per token, 47.17 tokens per second) 1.04.997.175 I slot print_timing: id 1 | task 189 | eval time = 8742.76 ms / 128 tokens ( 68.30 ms per token, 14.64 tokens per second) 1.04.997.176 I slot print_timing: id 1 | task 189 | total time = 8827.56 ms / 132 tokens 1.04.997.177 I slot print_timing: id 1 | task 189 | graphs reused = 1 1.04.997.179 I slot print_timing: id 1 | task 189 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.04.997.191 I statistics draft-mtp: #calls(b,g,a) = 58 239 3677, #gen drafts = 3692, #acc drafts = 2538, #gen tokens = 3692, #acc tokens = 2538, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.040, 399.584, 4.333 ms 1.04.997.218 I slot release: id 1 | task 189 | stop processing: n_tokens = 156, truncated = 0 1.05.000.666 I slot print_timing: id 9 | task 231 | n_decoded = 100, tg = 18.49 t/s, tg_3s = 18.49 t/s 1.05.006.253 I srv operator(): Chat format: peg-native 1.05.087.580 I slot print_timing: id 10 | task 212 | prompt eval time = 108.04 ms / 4 tokens ( 27.01 ms per token, 37.02 tokens per second) 1.05.087.583 I slot print_timing: id 10 | task 212 | eval time = 7187.38 ms / 128 tokens ( 56.15 ms per token, 17.81 tokens per second) 1.05.087.583 I slot print_timing: id 10 | task 212 | total time = 7295.42 ms / 132 tokens 1.05.087.584 I slot print_timing: id 10 | task 212 | graphs reused = 1 1.05.087.587 I slot print_timing: id 10 | task 212 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.05.087.597 I statistics draft-mtp: #calls(b,g,a) = 58 240 3692, #gen drafts = 3706, #acc drafts = 2547, #gen tokens = 3706, #acc tokens = 2547, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.040, 402.246, 4.350 ms 1.05.087.623 I slot release: id 10 | task 212 | stop processing: n_tokens = 155, truncated = 0 1.05.093.307 I slot get_availabl: id 1 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.05.093.309 I srv get_availabl: updating prompt cache 1.05.093.353 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.05.093.355 I srv alloc: - prompt is already in the cache, skipping 1.05.093.356 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.05.093.359 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.05.093.361 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.05.093.361 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.05.093.362 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.05.093.362 I srv get_availabl: prompt cache update took 0.05 ms 1.05.093.437 I slot launch_slot_: id 1 | task 300 | processing task, is_child = 0 1.05.093.438 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 1.05.093.455 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.05.093.456 I srv alloc: - prompt is already in the cache, skipping 1.05.096.337 I slot operator(): id 1 | task 300 | Checking checkpoint with [24, 24] against 28... 1.05.097.693 I srv operator(): Chat format: peg-native 1.05.100.255 W slot operator(): id 1 | task 300 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.05.100.296 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.05.184.873 I slot print_timing: id 7 | task 192 | prompt eval time = 88.58 ms / 4 tokens ( 22.14 ms per token, 45.16 tokens per second) 1.05.184.876 I slot print_timing: id 7 | task 192 | eval time = 8834.07 ms / 128 tokens ( 69.02 ms per token, 14.49 tokens per second) 1.05.184.877 I slot print_timing: id 7 | task 192 | total time = 8922.65 ms / 132 tokens 1.05.184.877 I slot print_timing: id 7 | task 192 | graphs reused = 1 1.05.184.880 I slot print_timing: id 7 | task 192 | draft acceptance = 0.41573 ( 37 accepted / 89 generated), mean acceptance length = 1.42, acceptance rate per position = (0.416) 1.05.184.890 I statistics draft-mtp: #calls(b,g,a) = 59 241 3706, #gen drafts = 3719, #acc drafts = 2557, #gen tokens = 3719, #acc tokens = 2557, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.041, 405.091, 4.366 ms 1.05.184.915 I slot release: id 7 | task 192 | stop processing: n_tokens = 156, truncated = 0 1.05.190.284 I slot get_availabl: id 10 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.05.190.285 I srv get_availabl: updating prompt cache 1.05.190.328 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.05.190.330 I srv alloc: - prompt is already in the cache, skipping 1.05.190.331 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.05.190.335 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.05.190.335 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.05.190.336 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.05.190.336 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.05.190.337 I srv get_availabl: prompt cache update took 0.05 ms 1.05.190.399 I slot launch_slot_: id 10 | task 302 | processing task, is_child = 0 1.05.190.407 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 1.05.190.423 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.05.190.424 I srv alloc: - prompt is already in the cache, skipping 1.05.193.033 I slot operator(): id 10 | task 302 | Checking checkpoint with [23, 23] against 27... 1.05.193.933 I srv operator(): Chat format: peg-native 1.05.196.936 W slot operator(): id 10 | task 302 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.05.196.979 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.05.287.407 I slot get_availabl: id 7 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.05.287.409 I srv get_availabl: updating prompt cache 1.05.287.452 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.05.287.454 I srv alloc: - prompt is already in the cache, skipping 1.05.287.455 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.05.287.457 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.05.287.459 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.05.287.459 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.05.287.460 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.05.287.460 I srv get_availabl: prompt cache update took 0.05 ms 1.05.287.523 I slot launch_slot_: id 7 | task 304 | processing task, is_child = 0 1.05.289.916 I slot operator(): id 7 | task 304 | Checking checkpoint with [24, 24] against 3... 1.05.289.918 W slot operator(): id 7 | task 304 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.05.289.921 W slot operator(): id 7 | task 304 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.05.291.449 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.05.412.897 I slot create_check: id 7 | task 304 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.05.505.111 I slot print_timing: id 3 | task 234 | prompt eval time = 212.50 ms / 25 tokens ( 8.50 ms per token, 117.65 tokens per second) 1.05.505.115 I slot print_timing: id 3 | task 234 | eval time = 5691.18 ms / 114 tokens ( 49.92 ms per token, 20.03 tokens per second) 1.05.505.115 I slot print_timing: id 3 | task 234 | total time = 5903.68 ms / 139 tokens 1.05.505.116 I slot print_timing: id 3 | task 234 | graphs reused = 1 1.05.505.118 I slot print_timing: id 3 | task 234 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.05.505.131 I statistics draft-mtp: #calls(b,g,a) = 61 244 3752, #gen drafts = 3763, #acc drafts = 2590, #gen tokens = 3763, #acc tokens = 2590, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.043, 412.382, 4.430 ms 1.05.505.161 I slot release: id 3 | task 234 | stop processing: n_tokens = 139, truncated = 0 1.05.513.957 I srv operator(): Chat format: peg-native 1.05.600.434 I slot get_availabl: id 3 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 1.05.600.436 I srv get_availabl: updating prompt cache 1.05.600.479 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.05.600.481 I srv alloc: - prompt is already in the cache, skipping 1.05.600.482 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 1.05.600.485 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.05.600.486 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.05.600.487 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.05.600.487 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.05.600.487 I srv get_availabl: prompt cache update took 0.05 ms 1.05.600.556 I slot launch_slot_: id 3 | task 308 | processing task, is_child = 0 1.05.602.584 I slot operator(): id 3 | task 308 | Checking checkpoint with [20, 20] against 3... 1.05.602.587 W slot operator(): id 3 | task 308 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.05.602.590 W slot operator(): id 3 | task 308 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.05.604.064 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.05.726.571 I slot create_check: id 3 | task 308 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.06.013.683 I slot print_timing: id 15 | task 227 | n_decoded = 101, tg = 15.02 t/s, tg_3s = 15.02 t/s 1.06.296.335 I slot print_timing: id 6 | task 247 | n_decoded = 100, tg = 18.59 t/s, tg_3s = 18.59 t/s 1.06.484.517 I slot print_timing: id 2 | task 210 | prompt eval time = 83.51 ms / 4 tokens ( 20.88 ms per token, 47.90 tokens per second) 1.06.484.521 I slot print_timing: id 2 | task 210 | eval time = 8700.41 ms / 128 tokens ( 67.97 ms per token, 14.71 tokens per second) 1.06.484.521 I slot print_timing: id 2 | task 210 | total time = 8783.92 ms / 132 tokens 1.06.484.522 I slot print_timing: id 2 | task 210 | graphs reused = 1 1.06.484.524 I slot print_timing: id 2 | task 210 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.06.484.541 I statistics draft-mtp: #calls(b,g,a) = 62 254 3904, #gen drafts = 3919, #acc drafts = 2694, #gen tokens = 3919, #acc tokens = 2694, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.045, 428.974, 4.624 ms 1.06.484.569 I slot release: id 2 | task 210 | stop processing: n_tokens = 156, truncated = 0 1.06.493.717 I srv operator(): Chat format: peg-native 1.06.581.054 I slot get_availabl: id 2 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 1.06.581.056 I srv get_availabl: updating prompt cache 1.06.581.098 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.06.581.101 I srv alloc: - prompt is already in the cache, skipping 1.06.581.102 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 1.06.581.105 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.06.581.105 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.06.581.106 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.06.581.106 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.06.581.107 I srv get_availabl: prompt cache update took 0.05 ms 1.06.581.171 I slot launch_slot_: id 2 | task 319 | processing task, is_child = 0 1.06.583.124 I slot operator(): id 2 | task 319 | Checking checkpoint with [24, 24] against 3... 1.06.583.127 W slot operator(): id 2 | task 319 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.06.583.128 W slot operator(): id 2 | task 319 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.06.584.861 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.06.707.143 I slot create_check: id 2 | task 319 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.06.796.437 I slot print_timing: id 9 | task 231 | prompt eval time = 207.17 ms / 28 tokens ( 7.40 ms per token, 135.16 tokens per second) 1.06.796.440 I slot print_timing: id 9 | task 231 | eval time = 7203.00 ms / 128 tokens ( 56.27 ms per token, 17.77 tokens per second) 1.06.796.440 I slot print_timing: id 9 | task 231 | total time = 7410.17 ms / 156 tokens 1.06.796.441 I slot print_timing: id 9 | task 231 | graphs reused = 1 1.06.796.444 I slot print_timing: id 9 | task 231 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.06.796.457 I statistics draft-mtp: #calls(b,g,a) = 63 257 3949, #gen drafts = 3963, #acc drafts = 2726, #gen tokens = 3963, #acc tokens = 2726, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.047, 435.327, 4.676 ms 1.06.796.487 I slot release: id 9 | task 231 | stop processing: n_tokens = 155, truncated = 0 1.06.805.428 I srv operator(): Chat format: peg-native 1.06.892.602 I slot get_availabl: id 9 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.06.892.605 I srv get_availabl: updating prompt cache 1.06.892.646 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.06.892.648 I srv alloc: - prompt is already in the cache, skipping 1.06.892.649 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.06.892.651 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.06.892.653 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.06.892.653 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.06.892.654 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.06.892.654 I srv get_availabl: prompt cache update took 0.05 ms 1.06.892.721 I slot launch_slot_: id 9 | task 323 | processing task, is_child = 0 1.06.895.237 I slot operator(): id 9 | task 323 | Checking checkpoint with [23, 23] against 3... 1.06.895.240 W slot operator(): id 9 | task 323 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.06.895.243 W slot operator(): id 9 | task 323 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.06.896.741 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.07.018.551 I slot create_check: id 9 | task 323 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.07.302.815 I slot print_timing: id 8 | task 263 | n_decoded = 101, tg = 20.02 t/s, tg_3s = 20.02 t/s 1.07.682.340 I slot print_timing: id 4 | task 246 | n_decoded = 101, tg = 14.93 t/s, tg_3s = 14.93 t/s 1.07.877.385 I slot print_timing: id 14 | task 266 | n_decoded = 100, tg = 18.58 t/s, tg_3s = 18.58 t/s 1.07.965.599 I slot print_timing: id 15 | task 227 | prompt eval time = 213.42 ms / 29 tokens ( 7.36 ms per token, 135.88 tokens per second) 1.07.965.602 I slot print_timing: id 15 | task 227 | eval time = 8678.33 ms / 128 tokens ( 67.80 ms per token, 14.75 tokens per second) 1.07.965.602 I slot print_timing: id 15 | task 227 | total time = 8891.75 ms / 157 tokens 1.07.965.603 I slot print_timing: id 15 | task 227 | graphs reused = 1 1.07.965.606 I slot print_timing: id 15 | task 227 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.07.965.620 I statistics draft-mtp: #calls(b,g,a) = 64 269 4136, #gen drafts = 4151, #acc drafts = 2852, #gen tokens = 4151, #acc tokens = 2852, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.048, 455.078, 4.907 ms 1.07.965.648 I slot release: id 15 | task 227 | stop processing: n_tokens = 156, truncated = 0 1.07.969.437 I slot print_timing: id 8 | task 263 | prompt eval time = 103.50 ms / 4 tokens ( 25.88 ms per token, 38.65 tokens per second) 1.07.969.439 I slot print_timing: id 8 | task 263 | eval time = 5711.50 ms / 114 tokens ( 50.10 ms per token, 19.96 tokens per second) 1.07.969.439 I slot print_timing: id 8 | task 263 | total time = 5815.00 ms / 118 tokens 1.07.969.440 I slot print_timing: id 8 | task 263 | graphs reused = 1 1.07.969.442 I slot print_timing: id 8 | task 263 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.07.969.449 I statistics draft-mtp: #calls(b,g,a) = 64 269 4145, #gen drafts = 4151, #acc drafts = 2859, #gen tokens = 4151, #acc tokens = 2859, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.048, 455.078, 4.920 ms 1.07.969.472 I slot release: id 8 | task 263 | stop processing: n_tokens = 139, truncated = 0 1.07.974.752 I srv operator(): Chat format: peg-native 1.07.977.685 I srv operator(): Chat format: peg-native 1.08.051.515 I slot print_timing: id 6 | task 247 | prompt eval time = 243.75 ms / 28 tokens ( 8.71 ms per token, 114.87 tokens per second) 1.08.051.518 I slot print_timing: id 6 | task 247 | eval time = 7135.19 ms / 128 tokens ( 55.74 ms per token, 17.94 tokens per second) 1.08.051.518 I slot print_timing: id 6 | task 247 | total time = 7378.94 ms / 156 tokens 1.08.051.519 I slot print_timing: id 6 | task 247 | graphs reused = 1 1.08.051.522 I slot print_timing: id 6 | task 247 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.08.051.534 I statistics draft-mtp: #calls(b,g,a) = 64 270 4151, #gen drafts = 4164, #acc drafts = 2863, #gen tokens = 4164, #acc tokens = 2863, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.048, 457.575, 4.927 ms 1.08.051.560 I slot release: id 6 | task 247 | stop processing: n_tokens = 155, truncated = 0 1.08.056.827 I slot get_availabl: id 15 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.08.056.829 I srv get_availabl: updating prompt cache 1.08.056.872 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.08.056.875 I srv alloc: - prompt is already in the cache, skipping 1.08.056.876 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.08.056.878 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.08.056.879 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.08.056.880 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.08.056.880 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.08.056.881 I srv get_availabl: prompt cache update took 0.05 ms 1.08.056.949 I slot launch_slot_: id 15 | task 336 | processing task, is_child = 0 1.08.056.950 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 1.08.056.967 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.08.056.969 I srv alloc: - prompt is already in the cache, skipping 1.08.056.969 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 1.08.056.984 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.08.056.986 I srv alloc: - prompt is already in the cache, skipping 1.08.056.988 I slot get_availabl: id 6 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.08.056.988 I srv get_availabl: updating prompt cache 1.08.057.001 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.08.057.002 I srv alloc: - prompt is already in the cache, skipping 1.08.057.003 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.08.057.003 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.08.057.004 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.08.057.004 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.08.057.006 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.08.057.006 I srv get_availabl: prompt cache update took 0.02 ms 1.08.057.023 I slot launch_slot_: id 6 | task 337 | processing task, is_child = 0 1.08.057.024 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 1.08.057.038 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.08.057.039 I srv alloc: - prompt is already in the cache, skipping 1.08.059.548 I slot operator(): id 6 | task 337 | Checking checkpoint with [23, 23] against 27... 1.08.060.582 I srv operator(): Chat format: peg-native 1.08.063.429 W slot operator(): id 6 | task 337 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.08.063.470 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.08.063.478 I slot operator(): id 15 | task 336 | Checking checkpoint with [24, 24] against 28... 1.08.066.988 W slot operator(): id 15 | task 336 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.08.067.015 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.08.159.530 I slot get_availabl: id 8 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.08.159.533 I srv get_availabl: updating prompt cache 1.08.159.573 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.08.159.575 I srv alloc: - prompt is already in the cache, skipping 1.08.159.576 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.08.159.579 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.08.159.580 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.08.159.581 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.08.159.582 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.08.159.582 I srv get_availabl: prompt cache update took 0.05 ms 1.08.159.649 I slot launch_slot_: id 8 | task 339 | processing task, is_child = 0 1.08.162.084 I slot operator(): id 8 | task 339 | Checking checkpoint with [20, 20] against 24... 1.08.165.997 W slot operator(): id 8 | task 339 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.08.166.032 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.08.733.395 I slot print_timing: id 0 | task 280 | n_decoded = 101, tg = 20.24 t/s, tg_3s = 20.24 t/s 1.08.929.415 I slot print_timing: id 13 | task 286 | n_decoded = 101, tg = 20.28 t/s, tg_3s = 20.28 t/s 1.09.120.725 I slot print_timing: id 12 | task 264 | n_decoded = 101, tg = 14.99 t/s, tg_3s = 14.99 t/s 1.09.215.620 I slot print_timing: id 11 | task 283 | n_decoded = 100, tg = 18.63 t/s, tg_3s = 18.63 t/s 1.09.402.231 I slot print_timing: id 0 | task 280 | prompt eval time = 199.77 ms / 25 tokens ( 7.99 ms per token, 125.15 tokens per second) 1.09.402.237 I slot print_timing: id 0 | task 280 | eval time = 5658.22 ms / 114 tokens ( 49.63 ms per token, 20.15 tokens per second) 1.09.402.237 I slot print_timing: id 0 | task 280 | total time = 5857.99 ms / 139 tokens 1.09.402.238 I slot print_timing: id 0 | task 280 | graphs reused = 1 1.09.402.241 I slot print_timing: id 0 | task 280 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.09.402.259 I statistics draft-mtp: #calls(b,g,a) = 67 284 4369, #gen drafts = 4384, #acc drafts = 3016, #gen tokens = 4384, #acc tokens = 3016, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.050, 480.023, 5.225 ms 1.09.402.285 I slot release: id 0 | task 280 | stop processing: n_tokens = 139, truncated = 0 1.09.412.839 I srv operator(): Chat format: peg-native 1.09.497.747 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 1.09.497.750 I srv get_availabl: updating prompt cache 1.09.497.791 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.09.497.793 I srv alloc: - prompt is already in the cache, skipping 1.09.497.794 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 1.09.497.797 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.09.497.798 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.09.497.799 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.09.497.801 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.09.497.801 I srv get_availabl: prompt cache update took 0.05 ms 1.09.497.868 I slot launch_slot_: id 0 | task 354 | processing task, is_child = 0 1.09.500.221 I slot operator(): id 0 | task 354 | Checking checkpoint with [20, 20] against 3... 1.09.500.223 W slot operator(): id 0 | task 354 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.09.500.225 W slot operator(): id 0 | task 354 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.09.501.961 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.09.600.935 I slot print_timing: id 4 | task 246 | prompt eval time = 245.17 ms / 29 tokens ( 8.45 ms per token, 118.29 tokens per second) 1.09.600.938 I slot print_timing: id 4 | task 246 | eval time = 8684.88 ms / 128 tokens ( 67.85 ms per token, 14.74 tokens per second) 1.09.600.938 I slot print_timing: id 4 | task 246 | total time = 8930.05 ms / 157 tokens 1.09.600.939 I slot print_timing: id 4 | task 246 | graphs reused = 1 1.09.600.942 I slot print_timing: id 4 | task 246 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.09.600.956 I statistics draft-mtp: #calls(b,g,a) = 67 286 4399, #gen drafts = 4412, #acc drafts = 3037, #gen tokens = 4412, #acc tokens = 3037, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.050, 483.780, 5.261 ms 1.09.600.981 I slot release: id 4 | task 246 | stop processing: n_tokens = 156, truncated = 0 1.09.601.196 I slot print_timing: id 14 | task 266 | prompt eval time = 215.26 ms / 28 tokens ( 7.69 ms per token, 130.07 tokens per second) 1.09.601.198 I slot print_timing: id 14 | task 266 | eval time = 7106.82 ms / 128 tokens ( 55.52 ms per token, 18.01 tokens per second) 1.09.601.199 I slot print_timing: id 14 | task 266 | total time = 7322.08 ms / 156 tokens 1.09.601.199 I slot print_timing: id 14 | task 266 | graphs reused = 1 1.09.601.200 I slot print_timing: id 14 | task 266 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.09.601.201 I statistics draft-mtp: #calls(b,g,a) = 67 286 4399, #gen drafts = 4412, #acc drafts = 3037, #gen tokens = 4412, #acc tokens = 3037, #mean acc len = 1.69, #acc rate/pos = (0.690), dur(b,g,a) = 0.050, 483.780, 5.261 ms 1.09.601.224 I slot release: id 14 | task 266 | stop processing: n_tokens = 155, truncated = 0 1.09.606.191 I slot print_timing: id 13 | task 286 | prompt eval time = 94.16 ms / 4 tokens ( 23.54 ms per token, 42.48 tokens per second) 1.09.606.193 I slot print_timing: id 13 | task 286 | eval time = 5657.40 ms / 114 tokens ( 49.63 ms per token, 20.15 tokens per second) 1.09.606.194 I slot print_timing: id 13 | task 286 | total time = 5751.56 ms / 118 tokens 1.09.606.194 I slot print_timing: id 13 | task 286 | graphs reused = 1 1.09.606.196 I slot print_timing: id 13 | task 286 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.09.606.202 I statistics draft-mtp: #calls(b,g,a) = 67 286 4411, #gen drafts = 4412, #acc drafts = 3046, #gen tokens = 4412, #acc tokens = 3046, #mean acc len = 1.69, #acc rate/pos = (0.691), dur(b,g,a) = 0.050, 483.780, 5.275 ms 1.09.606.224 I slot release: id 13 | task 286 | stop processing: n_tokens = 139, truncated = 0 1.09.610.072 I srv operator(): Chat format: peg-native 1.09.610.089 I srv operator(): Chat format: peg-native 1.09.615.153 I srv operator(): Chat format: peg-native 1.09.623.172 I slot create_check: id 0 | task 354 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.09.702.170 I slot get_availabl: id 14 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.09.702.172 I srv get_availabl: updating prompt cache 1.09.702.213 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.09.702.215 I srv alloc: - prompt is already in the cache, skipping 1.09.702.216 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.09.702.219 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.09.702.220 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.09.702.221 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.09.702.222 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.09.702.222 I srv get_availabl: prompt cache update took 0.05 ms 1.09.702.289 I slot launch_slot_: id 14 | task 358 | processing task, is_child = 0 1.09.702.290 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 1.09.702.308 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.09.702.309 I srv alloc: - prompt is already in the cache, skipping 1.09.702.309 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 1.09.702.324 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.09.702.326 I srv alloc: - prompt is already in the cache, skipping 1.09.702.327 I slot get_availabl: id 13 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.09.702.329 I srv get_availabl: updating prompt cache 1.09.702.342 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.09.702.343 I srv alloc: - prompt is already in the cache, skipping 1.09.702.344 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.09.702.344 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.09.702.345 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.09.702.346 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.09.702.346 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.09.702.346 I srv get_availabl: prompt cache update took 0.02 ms 1.09.702.364 I slot launch_slot_: id 13 | task 357 | processing task, is_child = 0 1.09.702.366 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 1.09.702.380 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.09.702.381 I srv alloc: - prompt is already in the cache, skipping 1.09.702.383 I slot get_availabl: id 4 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.09.702.384 I srv get_availabl: updating prompt cache 1.09.702.398 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.09.702.402 I srv alloc: - prompt is already in the cache, skipping 1.09.702.403 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.09.702.403 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.09.702.403 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.09.702.405 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.09.702.405 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.09.702.406 I srv get_availabl: prompt cache update took 0.02 ms 1.09.702.420 I slot launch_slot_: id 4 | task 359 | processing task, is_child = 0 1.09.705.107 I slot operator(): id 4 | task 359 | Checking checkpoint with [24, 24] against 28... 1.09.708.978 W slot operator(): id 4 | task 359 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.09.709.017 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.09.709.028 I slot operator(): id 13 | task 357 | Checking checkpoint with [20, 20] against 24... 1.09.712.573 W slot operator(): id 13 | task 357 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.09.712.609 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.09.712.615 I slot operator(): id 14 | task 358 | Checking checkpoint with [23, 23] against 27... 1.09.716.120 W slot operator(): id 14 | task 358 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.09.716.159 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.10.479.919 I slot print_timing: id 5 | task 284 | n_decoded = 101, tg = 15.23 t/s, tg_3s = 15.23 t/s 1.10.480.672 I slot print_timing: id 7 | task 304 | n_decoded = 101, tg = 20.29 t/s, tg_3s = 20.29 t/s 1.10.672.473 I slot print_timing: id 10 | task 302 | n_decoded = 100, tg = 18.55 t/s, tg_3s = 18.55 t/s 1.10.954.460 I slot print_timing: id 11 | task 283 | prompt eval time = 91.27 ms / 4 tokens ( 22.82 ms per token, 43.83 tokens per second) 1.10.954.464 I slot print_timing: id 11 | task 283 | eval time = 7107.01 ms / 128 tokens ( 55.52 ms per token, 18.01 tokens per second) 1.10.954.464 I slot print_timing: id 11 | task 283 | total time = 7198.28 ms / 132 tokens 1.10.954.465 I slot print_timing: id 11 | task 283 | graphs reused = 1 1.10.954.468 I slot print_timing: id 11 | task 283 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.10.954.484 I statistics draft-mtp: #calls(b,g,a) = 71 300 4613, #gen drafts = 4628, #acc drafts = 3179, #gen tokens = 4628, #acc tokens = 3179, #mean acc len = 1.69, #acc rate/pos = (0.689), dur(b,g,a) = 0.051, 507.083, 5.527 ms 1.10.954.512 I slot release: id 11 | task 283 | stop processing: n_tokens = 155, truncated = 0 1.10.963.867 I srv operator(): Chat format: peg-native 1.11.044.286 I slot print_timing: id 12 | task 264 | prompt eval time = 224.56 ms / 29 tokens ( 7.74 ms per token, 129.14 tokens per second) 1.11.044.289 I slot print_timing: id 12 | task 264 | eval time = 8661.37 ms / 128 tokens ( 67.67 ms per token, 14.78 tokens per second) 1.11.044.289 I slot print_timing: id 12 | task 264 | total time = 8885.93 ms / 157 tokens 1.11.044.290 I slot print_timing: id 12 | task 264 | graphs reused = 1 1.11.044.293 I slot print_timing: id 12 | task 264 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.11.044.305 I statistics draft-mtp: #calls(b,g,a) = 71 301 4628, #gen drafts = 4642, #acc drafts = 3190, #gen tokens = 4642, #acc tokens = 3190, #mean acc len = 1.69, #acc rate/pos = (0.689), dur(b,g,a) = 0.051, 509.012, 5.544 ms 1.11.044.332 I slot release: id 12 | task 264 | stop processing: n_tokens = 156, truncated = 0 1.11.050.203 I slot get_availabl: id 11 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.11.050.205 I srv get_availabl: updating prompt cache 1.11.050.247 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.11.050.249 I srv alloc: - prompt is already in the cache, skipping 1.11.050.250 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.11.050.253 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.11.050.254 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.11.050.255 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.11.050.255 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.11.050.255 I srv get_availabl: prompt cache update took 0.05 ms 1.11.050.321 I slot launch_slot_: id 11 | task 374 | processing task, is_child = 0 1.11.050.322 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 1.11.050.338 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.11.050.338 I srv alloc: - prompt is already in the cache, skipping 1.11.052.290 I slot operator(): id 11 | task 374 | Checking checkpoint with [23, 23] against 27... 1.11.052.806 I srv operator(): Chat format: peg-native 1.11.056.147 W slot operator(): id 11 | task 374 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.11.056.180 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.11.143.804 I slot print_timing: id 7 | task 304 | prompt eval time = 213.24 ms / 25 tokens ( 8.53 ms per token, 117.24 tokens per second) 1.11.143.807 I slot print_timing: id 7 | task 304 | eval time = 5640.62 ms / 114 tokens ( 49.48 ms per token, 20.21 tokens per second) 1.11.143.807 I slot print_timing: id 7 | task 304 | total time = 5853.87 ms / 139 tokens 1.11.143.808 I slot print_timing: id 7 | task 304 | graphs reused = 1 1.11.143.811 I slot print_timing: id 7 | task 304 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.11.143.827 I statistics draft-mtp: #calls(b,g,a) = 72 302 4650, #gen drafts = 4656, #acc drafts = 3204, #gen tokens = 4656, #acc tokens = 3204, #mean acc len = 1.69, #acc rate/pos = (0.689), dur(b,g,a) = 0.051, 510.924, 5.571 ms 1.11.143.859 I slot release: id 7 | task 304 | stop processing: n_tokens = 139, truncated = 0 1.11.146.387 I slot get_availabl: id 7 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.11.146.389 I srv get_availabl: updating prompt cache 1.11.146.438 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.11.146.440 I srv alloc: - prompt is already in the cache, skipping 1.11.146.441 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.11.146.444 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.11.146.446 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.11.146.446 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.11.146.447 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.11.146.447 I srv get_availabl: prompt cache update took 0.06 ms 1.11.146.511 I slot launch_slot_: id 7 | task 376 | processing task, is_child = 0 1.11.146.512 I slot process_sing: id 12 | task -1 | saving idle slot to prompt cache 1.11.146.529 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.11.146.530 I srv alloc: - prompt is already in the cache, skipping 1.11.149.221 I slot operator(): id 7 | task 376 | Checking checkpoint with [20, 20] against 24... 1.11.152.722 I srv operator(): Chat format: peg-native 1.11.153.177 W slot operator(): id 7 | task 376 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.11.153.215 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.11.243.381 I slot get_availabl: id 12 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.11.243.384 I srv get_availabl: updating prompt cache 1.11.243.432 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.11.243.434 I srv alloc: - prompt is already in the cache, skipping 1.11.243.435 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.11.243.438 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.11.243.438 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.11.243.440 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.11.243.440 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.11.243.441 I srv get_availabl: prompt cache update took 0.06 ms 1.11.243.507 I slot launch_slot_: id 12 | task 378 | processing task, is_child = 0 1.11.245.486 I slot operator(): id 12 | task 378 | Checking checkpoint with [24, 24] against 28... 1.11.249.408 W slot operator(): id 12 | task 378 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.11.249.444 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.11.910.631 I slot print_timing: id 1 | task 300 | n_decoded = 101, tg = 15.02 t/s, tg_3s = 15.02 t/s 1.12.009.029 I slot print_timing: id 9 | task 323 | n_decoded = 101, tg = 20.61 t/s, tg_3s = 20.61 t/s 1.12.102.045 I slot print_timing: id 2 | task 319 | n_decoded = 100, tg = 18.85 t/s, tg_3s = 18.85 t/s 1.12.387.144 I slot print_timing: id 5 | task 284 | prompt eval time = 94.97 ms / 4 tokens ( 23.74 ms per token, 42.12 tokens per second) 1.12.387.147 I slot print_timing: id 5 | task 284 | eval time = 8539.97 ms / 128 tokens ( 66.72 ms per token, 14.99 tokens per second) 1.12.387.148 I slot print_timing: id 5 | task 284 | total time = 8634.95 ms / 132 tokens 1.12.387.149 I slot print_timing: id 5 | task 284 | graphs reused = 1 1.12.387.151 I slot print_timing: id 5 | task 284 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.12.387.164 I statistics draft-mtp: #calls(b,g,a) = 74 315 4845, #gen drafts = 4859, #acc drafts = 3334, #gen tokens = 4859, #acc tokens = 3334, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.054, 532.730, 5.802 ms 1.12.387.193 I slot release: id 5 | task 284 | stop processing: n_tokens = 156, truncated = 0 1.12.387.417 I slot print_timing: id 10 | task 302 | prompt eval time = 88.79 ms / 4 tokens ( 22.20 ms per token, 45.05 tokens per second) 1.12.387.419 I slot print_timing: id 10 | task 302 | eval time = 7105.58 ms / 128 tokens ( 55.51 ms per token, 18.01 tokens per second) 1.12.387.419 I slot print_timing: id 10 | task 302 | total time = 7194.37 ms / 132 tokens 1.12.387.420 I slot print_timing: id 10 | task 302 | graphs reused = 1 1.12.387.421 I slot print_timing: id 10 | task 302 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.12.387.423 I statistics draft-mtp: #calls(b,g,a) = 74 315 4845, #gen drafts = 4859, #acc drafts = 3334, #gen tokens = 4859, #acc tokens = 3334, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.054, 532.730, 5.802 ms 1.12.387.441 I slot release: id 10 | task 302 | stop processing: n_tokens = 155, truncated = 0 1.12.389.163 I slot print_timing: id 3 | task 308 | n_decoded = 101, tg = 15.37 t/s, tg_3s = 15.37 t/s 1.12.396.830 I srv operator(): Chat format: peg-native 1.12.396.860 I srv operator(): Chat format: peg-native 1.12.479.560 I slot get_availabl: id 10 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.12.479.562 I srv get_availabl: updating prompt cache 1.12.479.608 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.12.479.610 I srv alloc: - prompt is already in the cache, skipping 1.12.479.611 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.12.479.614 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.12.479.615 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.12.479.616 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.12.479.617 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.12.479.618 I srv get_availabl: prompt cache update took 0.06 ms 1.12.479.686 I slot launch_slot_: id 10 | task 392 | processing task, is_child = 0 1.12.479.688 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 1.12.479.706 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.12.479.707 I srv alloc: - prompt is already in the cache, skipping 1.12.479.709 I slot get_availabl: id 5 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.12.479.709 I srv get_availabl: updating prompt cache 1.12.479.723 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.12.479.724 I srv alloc: - prompt is already in the cache, skipping 1.12.479.724 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.12.479.725 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.12.479.726 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.12.479.726 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.12.479.726 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.12.479.727 I srv get_availabl: prompt cache update took 0.02 ms 1.12.479.745 I slot launch_slot_: id 5 | task 393 | processing task, is_child = 0 1.12.482.433 I slot operator(): id 5 | task 393 | Checking checkpoint with [24, 24] against 3... 1.12.482.436 W slot operator(): id 5 | task 393 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.12.482.438 W slot operator(): id 5 | task 393 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.12.484.355 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.12.484.385 I slot operator(): id 10 | task 392 | Checking checkpoint with [23, 23] against 27... 1.12.488.221 W slot operator(): id 10 | task 392 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.12.488.249 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.12.612.080 I slot create_check: id 5 | task 393 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.12.706.246 I slot print_timing: id 9 | task 323 | prompt eval time = 213.45 ms / 25 tokens ( 8.54 ms per token, 117.13 tokens per second) 1.12.706.249 I slot print_timing: id 9 | task 323 | eval time = 5597.54 ms / 114 tokens ( 49.10 ms per token, 20.37 tokens per second) 1.12.706.250 I slot print_timing: id 9 | task 323 | total time = 5810.99 ms / 139 tokens 1.12.706.251 I slot print_timing: id 9 | task 323 | graphs reused = 1 1.12.706.253 I slot print_timing: id 9 | task 323 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.12.706.274 I statistics draft-mtp: #calls(b,g,a) = 76 318 4896, #gen drafts = 4902, #acc drafts = 3371, #gen tokens = 4902, #acc tokens = 3371, #mean acc len = 1.69, #acc rate/pos = (0.689), dur(b,g,a) = 0.055, 540.103, 5.867 ms 1.12.706.303 I slot release: id 9 | task 323 | stop processing: n_tokens = 139, truncated = 0 1.12.716.555 I srv operator(): Chat format: peg-native 1.12.799.358 I slot get_availabl: id 9 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 1.12.799.361 I srv get_availabl: updating prompt cache 1.12.799.411 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.12.799.413 I srv alloc: - prompt is already in the cache, skipping 1.12.799.414 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 1.12.799.417 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.12.799.418 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.12.799.419 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.12.799.419 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.12.799.420 I srv get_availabl: prompt cache update took 0.06 ms 1.12.799.487 I slot launch_slot_: id 9 | task 397 | processing task, is_child = 0 1.12.801.794 I slot operator(): id 9 | task 397 | Checking checkpoint with [20, 20] against 3... 1.12.801.797 W slot operator(): id 9 | task 397 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.12.801.800 W slot operator(): id 9 | task 397 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.12.803.448 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.12.926.790 I slot create_check: id 9 | task 397 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.13.209.555 I slot print_timing: id 8 | task 339 | n_decoded = 101, tg = 20.39 t/s, tg_3s = 20.39 t/s 1.13.495.056 I slot print_timing: id 6 | task 337 | n_decoded = 100, tg = 18.72 t/s, tg_3s = 18.72 t/s 1.13.873.526 I slot print_timing: id 1 | task 300 | prompt eval time = 88.31 ms / 4 tokens ( 22.08 ms per token, 45.29 tokens per second) 1.13.873.529 I slot print_timing: id 1 | task 300 | eval time = 8688.85 ms / 128 tokens ( 67.88 ms per token, 14.73 tokens per second) 1.13.873.529 I slot print_timing: id 1 | task 300 | total time = 8777.16 ms / 132 tokens 1.13.873.530 I slot print_timing: id 1 | task 300 | graphs reused = 1 1.13.873.532 I slot print_timing: id 1 | task 300 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.13.873.548 I statistics draft-mtp: #calls(b,g,a) = 77 330 5075, #gen drafts = 5089, #acc drafts = 3492, #gen tokens = 5089, #acc tokens = 3492, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.056, 560.064, 6.089 ms 1.13.873.576 I slot release: id 1 | task 300 | stop processing: n_tokens = 156, truncated = 0 1.13.873.789 I slot print_timing: id 2 | task 319 | prompt eval time = 213.09 ms / 28 tokens ( 7.61 ms per token, 131.40 tokens per second) 1.13.873.791 I slot print_timing: id 2 | task 319 | eval time = 7077.57 ms / 128 tokens ( 55.29 ms per token, 18.09 tokens per second) 1.13.873.791 I slot print_timing: id 2 | task 319 | total time = 7290.65 ms / 156 tokens 1.13.873.791 I slot print_timing: id 2 | task 319 | graphs reused = 1 1.13.873.792 I slot print_timing: id 2 | task 319 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.13.873.794 I statistics draft-mtp: #calls(b,g,a) = 77 330 5075, #gen drafts = 5089, #acc drafts = 3492, #gen tokens = 5089, #acc tokens = 3492, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.056, 560.064, 6.089 ms 1.13.873.824 I slot release: id 2 | task 319 | stop processing: n_tokens = 155, truncated = 0 1.13.876.870 I slot print_timing: id 8 | task 339 | prompt eval time = 94.21 ms / 4 tokens ( 23.55 ms per token, 42.46 tokens per second) 1.13.876.873 I slot print_timing: id 8 | task 339 | eval time = 5620.56 ms / 114 tokens ( 49.30 ms per token, 20.28 tokens per second) 1.13.876.873 I slot print_timing: id 8 | task 339 | total time = 5714.77 ms / 118 tokens 1.13.876.874 I slot print_timing: id 8 | task 339 | graphs reused = 1 1.13.876.876 I slot print_timing: id 8 | task 339 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.13.876.884 I statistics draft-mtp: #calls(b,g,a) = 77 330 5082, #gen drafts = 5089, #acc drafts = 3498, #gen tokens = 5089, #acc tokens = 3498, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.056, 560.064, 6.099 ms 1.13.876.909 I slot release: id 8 | task 339 | stop processing: n_tokens = 139, truncated = 0 1.13.884.199 I srv operator(): Chat format: peg-native 1.13.884.209 I srv operator(): Chat format: peg-native 1.13.885.192 I srv operator(): Chat format: peg-native 1.13.959.358 I slot get_availabl: id 8 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.13.959.360 I srv get_availabl: updating prompt cache 1.13.959.407 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.13.959.410 I srv alloc: - prompt is already in the cache, skipping 1.13.959.411 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.13.959.413 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.13.959.414 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.13.959.415 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.13.959.415 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.13.959.416 I srv get_availabl: prompt cache update took 0.06 ms 1.13.959.482 I slot launch_slot_: id 8 | task 411 | processing task, is_child = 0 1.13.959.483 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 1.13.959.500 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.13.959.501 I srv alloc: - prompt is already in the cache, skipping 1.13.959.501 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 1.13.959.516 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.13.959.517 I srv alloc: - prompt is already in the cache, skipping 1.13.959.519 I slot get_availabl: id 2 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.13.959.519 I srv get_availabl: updating prompt cache 1.13.959.534 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.13.959.535 I srv alloc: - prompt is already in the cache, skipping 1.13.959.535 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.13.959.536 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.13.959.536 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.13.959.537 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.13.959.537 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.13.959.537 I srv get_availabl: prompt cache update took 0.02 ms 1.13.959.554 I slot launch_slot_: id 2 | task 410 | processing task, is_child = 0 1.13.959.555 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache 1.13.959.569 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.13.959.572 I srv alloc: - prompt is already in the cache, skipping 1.13.959.573 I slot get_availabl: id 1 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.13.959.574 I srv get_availabl: updating prompt cache 1.13.959.587 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.13.959.587 I srv alloc: - prompt is already in the cache, skipping 1.13.959.588 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.13.959.588 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.13.959.589 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.13.959.589 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.13.959.589 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.13.959.592 I srv get_availabl: prompt cache update took 0.02 ms 1.13.959.607 I slot launch_slot_: id 1 | task 412 | processing task, is_child = 0 1.13.962.136 I slot operator(): id 1 | task 412 | Checking checkpoint with [24, 24] against 28... 1.13.966.042 W slot operator(): id 1 | task 412 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.13.966.078 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.13.966.097 I slot operator(): id 2 | task 410 | Checking checkpoint with [23, 23] against 27... 1.13.969.663 W slot operator(): id 2 | task 410 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.13.969.696 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.13.969.710 I slot operator(): id 8 | task 411 | Checking checkpoint with [20, 20] against 24... 1.13.973.218 W slot operator(): id 8 | task 411 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.13.973.241 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.14.352.931 I slot print_timing: id 3 | task 308 | prompt eval time = 214.09 ms / 29 tokens ( 7.38 ms per token, 135.46 tokens per second) 1.14.352.934 I slot print_timing: id 3 | task 308 | eval time = 8536.23 ms / 128 tokens ( 66.69 ms per token, 14.99 tokens per second) 1.14.352.935 I slot print_timing: id 3 | task 308 | total time = 8750.32 ms / 157 tokens 1.14.352.936 I slot print_timing: id 3 | task 308 | graphs reused = 1 1.14.352.938 I slot print_timing: id 3 | task 308 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.14.352.952 I statistics draft-mtp: #calls(b,g,a) = 80 335 5147, #gen drafts = 5162, #acc drafts = 3540, #gen tokens = 5162, #acc tokens = 3540, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.058, 569.993, 6.180 ms 1.14.352.979 I slot release: id 3 | task 308 | stop processing: n_tokens = 156, truncated = 0 1.14.361.387 I srv operator(): Chat format: peg-native 1.14.450.080 I slot get_availabl: id 3 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 1.14.450.083 I srv get_availabl: updating prompt cache 1.14.450.125 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.14.450.127 I srv alloc: - prompt is already in the cache, skipping 1.14.450.128 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 1.14.450.131 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.14.450.132 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.14.450.133 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.14.450.133 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.14.450.134 I srv get_availabl: prompt cache update took 0.05 ms 1.14.450.203 I slot launch_slot_: id 3 | task 418 | processing task, is_child = 0 1.14.452.262 I slot operator(): id 3 | task 418 | Checking checkpoint with [24, 24] against 3... 1.14.452.265 W slot operator(): id 3 | task 418 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.14.452.268 W slot operator(): id 3 | task 418 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.14.454.053 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.14.575.717 I slot create_check: id 3 | task 418 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.14.766.020 I slot print_timing: id 13 | task 357 | n_decoded = 101, tg = 20.38 t/s, tg_3s = 20.38 t/s 1.14.862.258 I slot print_timing: id 15 | task 336 | n_decoded = 101, tg = 15.06 t/s, tg_3s = 15.06 t/s 1.15.147.415 I slot print_timing: id 14 | task 358 | n_decoded = 100, tg = 18.74 t/s, tg_3s = 18.74 t/s 1.15.237.187 I slot print_timing: id 6 | task 337 | prompt eval time = 94.10 ms / 4 tokens ( 23.52 ms per token, 42.51 tokens per second) 1.15.237.190 I slot print_timing: id 6 | task 337 | eval time = 7083.51 ms / 128 tokens ( 55.34 ms per token, 18.07 tokens per second) 1.15.237.191 I slot print_timing: id 6 | task 337 | total time = 7177.61 ms / 132 tokens 1.15.237.191 I slot print_timing: id 6 | task 337 | graphs reused = 1 1.15.237.194 I slot print_timing: id 6 | task 337 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.15.237.210 I statistics draft-mtp: #calls(b,g,a) = 81 344 5287, #gen drafts = 5302, #acc drafts = 3635, #gen tokens = 5302, #acc tokens = 3635, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.060, 585.426, 6.352 ms 1.15.237.238 I slot release: id 6 | task 337 | stop processing: n_tokens = 155, truncated = 0 1.15.246.270 I srv operator(): Chat format: peg-native 1.15.334.235 I slot get_availabl: id 6 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.15.334.237 I srv get_availabl: updating prompt cache 1.15.334.281 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.15.334.283 I srv alloc: - prompt is already in the cache, skipping 1.15.334.284 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.15.334.287 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.15.334.287 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.15.334.288 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.15.334.288 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.15.334.289 I srv get_availabl: prompt cache update took 0.05 ms 1.15.334.355 I slot launch_slot_: id 6 | task 428 | processing task, is_child = 0 1.15.336.708 I slot operator(): id 6 | task 428 | Checking checkpoint with [23, 23] against 3... 1.15.336.710 W slot operator(): id 6 | task 428 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.15.336.712 W slot operator(): id 6 | task 428 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.15.339.023 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.15.444.245 I slot print_timing: id 13 | task 357 | prompt eval time = 101.66 ms / 4 tokens ( 25.41 ms per token, 39.35 tokens per second) 1.15.444.248 I slot print_timing: id 13 | task 357 | eval time = 5633.53 ms / 114 tokens ( 49.42 ms per token, 20.24 tokens per second) 1.15.444.249 I slot print_timing: id 13 | task 357 | total time = 5735.19 ms / 118 tokens 1.15.444.250 I slot print_timing: id 13 | task 357 | graphs reused = 1 1.15.444.253 I slot print_timing: id 13 | task 357 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.15.444.264 I statistics draft-mtp: #calls(b,g,a) = 81 346 5330, #gen drafts = 5332, #acc drafts = 3668, #gen tokens = 5332, #acc tokens = 3668, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.060, 589.846, 6.404 ms 1.15.444.293 I slot release: id 13 | task 357 | stop processing: n_tokens = 139, truncated = 0 1.15.452.711 I srv operator(): Chat format: peg-native 1.15.461.969 I slot create_check: id 6 | task 428 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.15.552.347 I slot get_availabl: id 13 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 1.15.552.350 I srv get_availabl: updating prompt cache 1.15.552.392 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.15.552.394 I srv alloc: - prompt is already in the cache, skipping 1.15.552.395 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 1.15.552.398 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.15.552.399 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.15.552.408 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.15.552.408 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.15.552.408 I srv get_availabl: prompt cache update took 0.06 ms 1.15.552.475 I slot launch_slot_: id 13 | task 431 | processing task, is_child = 0 1.15.554.398 I slot operator(): id 13 | task 431 | Checking checkpoint with [20, 20] against 3... 1.15.554.405 W slot operator(): id 13 | task 431 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.15.554.406 W slot operator(): id 13 | task 431 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.15.556.360 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.15.678.607 I slot create_check: id 13 | task 431 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.16.248.461 I slot print_timing: id 7 | task 376 | n_decoded = 101, tg = 20.16 t/s, tg_3s = 20.16 t/s 1.16.436.195 I slot print_timing: id 0 | task 354 | n_decoded = 101, tg = 14.99 t/s, tg_3s = 14.99 t/s 1.16.437.830 I slot print_timing: id 4 | task 359 | n_decoded = 101, tg = 15.24 t/s, tg_3s = 15.24 t/s 1.16.536.080 I slot print_timing: id 11 | task 374 | n_decoded = 100, tg = 18.54 t/s, tg_3s = 18.54 t/s 1.16.815.917 I slot print_timing: id 15 | task 336 | prompt eval time = 90.44 ms / 4 tokens ( 22.61 ms per token, 44.23 tokens per second) 1.16.815.920 I slot print_timing: id 15 | task 336 | eval time = 8661.97 ms / 128 tokens ( 67.67 ms per token, 14.78 tokens per second) 1.16.815.921 I slot print_timing: id 15 | task 336 | total time = 8752.41 ms / 132 tokens 1.16.815.921 I slot print_timing: id 15 | task 336 | graphs reused = 1 1.16.815.923 I slot print_timing: id 15 | task 336 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.16.815.937 I statistics draft-mtp: #calls(b,g,a) = 83 360 5536, #gen drafts = 5551, #acc drafts = 3807, #gen tokens = 5551, #acc tokens = 3807, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.062, 612.153, 6.659 ms 1.16.815.962 I slot release: id 15 | task 336 | stop processing: n_tokens = 156, truncated = 0 1.16.824.633 I srv operator(): Chat format: peg-native 1.16.904.337 I slot print_timing: id 14 | task 358 | prompt eval time = 98.34 ms / 4 tokens ( 24.58 ms per token, 40.68 tokens per second) 1.16.904.340 I slot print_timing: id 14 | task 358 | eval time = 7093.36 ms / 128 tokens ( 55.42 ms per token, 18.05 tokens per second) 1.16.904.341 I slot print_timing: id 14 | task 358 | total time = 7191.70 ms / 132 tokens 1.16.904.341 I slot print_timing: id 14 | task 358 | graphs reused = 1 1.16.904.344 I slot print_timing: id 14 | task 358 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.16.904.355 I statistics draft-mtp: #calls(b,g,a) = 83 361 5551, #gen drafts = 5565, #acc drafts = 3816, #gen tokens = 5565, #acc tokens = 3816, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.062, 613.521, 6.678 ms 1.16.904.381 I slot release: id 14 | task 358 | stop processing: n_tokens = 155, truncated = 0 1.16.907.652 I slot print_timing: id 7 | task 376 | prompt eval time = 88.53 ms / 4 tokens ( 22.13 ms per token, 45.18 tokens per second) 1.16.907.654 I slot print_timing: id 7 | task 376 | eval time = 5669.89 ms / 114 tokens ( 49.74 ms per token, 20.11 tokens per second) 1.16.907.655 I slot print_timing: id 7 | task 376 | total time = 5758.41 ms / 118 tokens 1.16.907.656 I slot print_timing: id 7 | task 376 | graphs reused = 1 1.16.907.658 I slot print_timing: id 7 | task 376 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.16.907.665 I statistics draft-mtp: #calls(b,g,a) = 83 361 5559, #gen drafts = 5565, #acc drafts = 3822, #gen tokens = 5565, #acc tokens = 3822, #mean acc len = 1.69, #acc rate/pos = (0.688), dur(b,g,a) = 0.062, 613.521, 6.687 ms 1.16.907.689 I slot release: id 7 | task 376 | stop processing: n_tokens = 139, truncated = 0 1.16.910.341 I slot get_availabl: id 14 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.16.910.342 I srv get_availabl: updating prompt cache 1.16.910.385 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.16.910.387 I srv alloc: - prompt is already in the cache, skipping 1.16.910.388 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.16.910.391 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.16.910.393 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.16.910.393 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.16.910.393 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.16.910.394 I srv get_availabl: prompt cache update took 0.05 ms 1.16.910.467 I slot launch_slot_: id 14 | task 446 | processing task, is_child = 0 1.16.910.468 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 1.16.910.485 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.16.910.486 I srv alloc: - prompt is already in the cache, skipping 1.16.910.486 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 1.16.910.501 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.16.910.503 I srv alloc: - prompt is already in the cache, skipping 1.16.912.657 I slot operator(): id 14 | task 446 | Checking checkpoint with [23, 23] against 27... 1.16.913.247 I srv operator(): Chat format: peg-native 1.16.916.445 I srv operator(): Chat format: peg-native 1.16.916.558 W slot operator(): id 14 | task 446 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.16.916.594 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.17.001.275 I slot get_availabl: id 7 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.17.001.278 I srv get_availabl: updating prompt cache 1.17.001.319 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.17.001.321 I srv alloc: - prompt is already in the cache, skipping 1.17.001.322 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.17.001.325 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.17.001.326 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.17.001.327 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.17.001.327 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.17.001.329 I srv get_availabl: prompt cache update took 0.05 ms 1.17.001.397 I slot launch_slot_: id 7 | task 448 | processing task, is_child = 0 1.17.001.399 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 1.17.001.423 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.17.001.424 I srv alloc: - prompt is already in the cache, skipping 1.17.001.426 I slot get_availabl: id 15 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.17.001.426 I srv get_availabl: updating prompt cache 1.17.001.440 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.17.001.441 I srv alloc: - prompt is already in the cache, skipping 1.17.001.441 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.17.001.442 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.17.001.442 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.17.001.443 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.17.001.443 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.17.001.443 I srv get_availabl: prompt cache update took 0.02 ms 1.17.001.463 I slot launch_slot_: id 15 | task 449 | processing task, is_child = 0 1.17.003.695 I slot operator(): id 7 | task 448 | Checking checkpoint with [20, 20] against 24... 1.17.007.557 W slot operator(): id 7 | task 448 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.17.007.587 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.17.007.606 I slot operator(): id 15 | task 449 | Checking checkpoint with [24, 24] against 28... 1.17.011.123 W slot operator(): id 15 | task 449 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.17.011.157 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.17.677.268 I slot print_timing: id 5 | task 393 | n_decoded = 101, tg = 20.30 t/s, tg_3s = 20.30 t/s 1.17.966.003 I slot print_timing: id 10 | task 392 | n_decoded = 100, tg = 18.60 t/s, tg_3s = 18.60 t/s 1.17.966.774 I slot print_timing: id 12 | task 378 | n_decoded = 101, tg = 15.24 t/s, tg_3s = 15.24 t/s 1.18.247.775 I slot print_timing: id 11 | task 374 | prompt eval time = 88.64 ms / 4 tokens ( 22.16 ms per token, 45.12 tokens per second) 1.18.247.778 I slot print_timing: id 11 | task 374 | eval time = 7106.82 ms / 128 tokens ( 55.52 ms per token, 18.01 tokens per second) 1.18.247.778 I slot print_timing: id 11 | task 374 | total time = 7195.46 ms / 132 tokens 1.18.247.779 I slot print_timing: id 11 | task 374 | graphs reused = 1 1.18.247.782 I slot print_timing: id 11 | task 374 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.18.247.795 I statistics draft-mtp: #calls(b,g,a) = 86 375 5768, #gen drafts = 5783, #acc drafts = 3960, #gen tokens = 5783, #acc tokens = 3960, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.063, 635.910, 6.934 ms 1.18.247.821 I slot release: id 11 | task 374 | stop processing: n_tokens = 155, truncated = 0 1.18.257.664 I srv operator(): Chat format: peg-native 1.18.337.344 I slot print_timing: id 0 | task 354 | prompt eval time = 196.99 ms / 29 tokens ( 6.79 ms per token, 147.22 tokens per second) 1.18.337.347 I slot print_timing: id 0 | task 354 | eval time = 8640.10 ms / 128 tokens ( 67.50 ms per token, 14.81 tokens per second) 1.18.337.348 I slot print_timing: id 0 | task 354 | total time = 8837.09 ms / 157 tokens 1.18.337.349 I slot print_timing: id 0 | task 354 | graphs reused = 1 1.18.337.351 I slot print_timing: id 0 | task 354 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.18.337.365 I statistics draft-mtp: #calls(b,g,a) = 86 376 5783, #gen drafts = 5796, #acc drafts = 3970, #gen tokens = 5796, #acc tokens = 3970, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.063, 638.469, 6.952 ms 1.18.337.391 I slot release: id 0 | task 354 | stop processing: n_tokens = 156, truncated = 0 1.18.337.604 I slot print_timing: id 4 | task 359 | prompt eval time = 105.31 ms / 4 tokens ( 26.33 ms per token, 37.98 tokens per second) 1.18.337.606 I slot print_timing: id 4 | task 359 | eval time = 8527.18 ms / 128 tokens ( 66.62 ms per token, 15.01 tokens per second) 1.18.337.606 I slot print_timing: id 4 | task 359 | total time = 8632.49 ms / 132 tokens 1.18.337.606 I slot print_timing: id 4 | task 359 | graphs reused = 1 1.18.337.607 I slot print_timing: id 4 | task 359 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.18.337.609 I statistics draft-mtp: #calls(b,g,a) = 86 376 5783, #gen drafts = 5796, #acc drafts = 3970, #gen tokens = 5796, #acc tokens = 3970, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.063, 638.469, 6.952 ms 1.18.337.626 I slot release: id 4 | task 359 | stop processing: n_tokens = 156, truncated = 0 1.18.339.120 I slot print_timing: id 5 | task 393 | prompt eval time = 220.25 ms / 25 tokens ( 8.81 ms per token, 113.51 tokens per second) 1.18.339.122 I slot print_timing: id 5 | task 393 | eval time = 5636.43 ms / 114 tokens ( 49.44 ms per token, 20.23 tokens per second) 1.18.339.123 I slot print_timing: id 5 | task 393 | total time = 5856.67 ms / 139 tokens 1.18.339.123 I slot print_timing: id 5 | task 393 | graphs reused = 1 1.18.339.125 I slot print_timing: id 5 | task 393 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.18.339.129 I statistics draft-mtp: #calls(b,g,a) = 86 376 5787, #gen drafts = 5796, #acc drafts = 3972, #gen tokens = 5796, #acc tokens = 3972, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.063, 638.469, 6.957 ms 1.18.339.149 I slot release: id 5 | task 393 | stop processing: n_tokens = 139, truncated = 0 1.18.342.806 I slot get_availabl: id 11 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.18.342.808 I srv get_availabl: updating prompt cache 1.18.342.849 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.18.342.852 I srv alloc: - prompt is already in the cache, skipping 1.18.342.852 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.18.342.855 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.18.342.856 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.18.342.856 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.18.342.857 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.18.342.858 I srv get_availabl: prompt cache update took 0.05 ms 1.18.342.927 I slot launch_slot_: id 11 | task 464 | processing task, is_child = 0 1.18.342.929 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 1.18.342.946 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.18.342.947 I srv alloc: - prompt is already in the cache, skipping 1.18.342.948 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 1.18.342.963 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.18.342.965 I srv alloc: - prompt is already in the cache, skipping 1.18.342.965 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 1.18.342.980 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.18.342.982 I srv alloc: - prompt is already in the cache, skipping 1.18.345.494 I slot operator(): id 11 | task 464 | Checking checkpoint with [23, 23] against 27... 1.18.347.887 I srv operator(): Chat format: peg-native 1.18.348.131 I srv operator(): Chat format: peg-native 1.18.348.398 I srv operator(): Chat format: peg-native 1.18.349.337 W slot operator(): id 11 | task 464 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.18.349.371 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.18.428.069 I slot get_availabl: id 5 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.18.428.072 I srv get_availabl: updating prompt cache 1.18.428.112 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.18.428.114 I srv alloc: - prompt is already in the cache, skipping 1.18.428.115 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.18.428.118 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.18.428.119 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.18.428.120 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.18.428.121 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.18.428.121 I srv get_availabl: prompt cache update took 0.05 ms 1.18.428.186 I slot launch_slot_: id 5 | task 466 | processing task, is_child = 0 1.18.428.188 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache 1.18.428.204 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.18.428.206 I srv alloc: - prompt is already in the cache, skipping 1.18.428.206 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 1.18.428.221 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.18.428.223 I srv alloc: - prompt is already in the cache, skipping 1.18.428.225 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.18.428.225 I srv get_availabl: updating prompt cache 1.18.428.238 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.18.428.240 I srv alloc: - prompt is already in the cache, skipping 1.18.428.240 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.18.428.240 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.18.428.241 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.18.428.241 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.18.428.241 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.18.428.242 I srv get_availabl: prompt cache update took 0.02 ms 1.18.428.261 I slot launch_slot_: id 0 | task 467 | processing task, is_child = 0 1.18.428.262 I slot process_sing: id 4 | task -1 | saving idle slot to prompt cache 1.18.428.276 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.18.428.277 I srv alloc: - prompt is already in the cache, skipping 1.18.428.279 I slot get_availabl: id 4 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 1.18.428.279 I srv get_availabl: updating prompt cache 1.18.428.293 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.18.428.294 I srv alloc: - prompt is already in the cache, skipping 1.18.428.294 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 1.18.428.295 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.18.428.295 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.18.428.297 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.18.428.297 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.18.428.297 I srv get_availabl: prompt cache update took 0.02 ms 1.18.428.314 I slot launch_slot_: id 4 | task 468 | processing task, is_child = 0 1.18.430.219 I slot operator(): id 0 | task 467 | Checking checkpoint with [24, 24] against 28... 1.18.434.102 W slot operator(): id 0 | task 467 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.18.434.139 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.18.434.150 I slot operator(): id 4 | task 468 | Checking checkpoint with [24, 24] against 3... 1.18.434.151 W slot operator(): id 4 | task 468 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.18.434.152 W slot operator(): id 4 | task 468 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.18.435.874 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.18.435.898 I slot operator(): id 5 | task 466 | Checking checkpoint with [20, 20] against 24... 1.18.439.516 W slot operator(): id 5 | task 466 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.18.439.544 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.18.565.551 I slot create_check: id 4 | task 468 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.19.041.269 I slot print_timing: id 8 | task 411 | n_decoded = 101, tg = 20.31 t/s, tg_3s = 20.31 t/s 1.19.419.685 I slot print_timing: id 2 | task 410 | n_decoded = 100, tg = 18.69 t/s, tg_3s = 18.69 t/s 1.19.614.070 I slot print_timing: id 9 | task 397 | n_decoded = 101, tg = 15.31 t/s, tg_3s = 15.31 t/s 1.19.705.730 I slot print_timing: id 10 | task 392 | prompt eval time = 105.92 ms / 4 tokens ( 26.48 ms per token, 37.77 tokens per second) 1.19.705.733 I slot print_timing: id 10 | task 392 | eval time = 7115.40 ms / 128 tokens ( 55.59 ms per token, 17.99 tokens per second) 1.19.705.733 I slot print_timing: id 10 | task 392 | total time = 7221.32 ms / 132 tokens 1.19.705.734 I slot print_timing: id 10 | task 392 | graphs reused = 1 1.19.705.737 I slot print_timing: id 10 | task 392 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.19.705.751 I statistics draft-mtp: #calls(b,g,a) = 90 390 5996, #gen drafts = 6011, #acc drafts = 4116, #gen tokens = 6011, #acc tokens = 4116, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.066, 661.638, 7.203 ms 1.19.705.777 I slot release: id 10 | task 392 | stop processing: n_tokens = 155, truncated = 0 1.19.709.848 I slot print_timing: id 8 | task 411 | prompt eval time = 98.83 ms / 4 tokens ( 24.71 ms per token, 40.47 tokens per second) 1.19.709.850 I slot print_timing: id 8 | task 411 | eval time = 5641.29 ms / 114 tokens ( 49.49 ms per token, 20.21 tokens per second) 1.19.709.850 I slot print_timing: id 8 | task 411 | total time = 5740.12 ms / 118 tokens 1.19.709.851 I slot print_timing: id 8 | task 411 | graphs reused = 1 1.19.709.853 I slot print_timing: id 8 | task 411 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.19.709.860 I statistics draft-mtp: #calls(b,g,a) = 90 390 6005, #gen drafts = 6011, #acc drafts = 4125, #gen tokens = 6011, #acc tokens = 4125, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.066, 661.638, 7.215 ms 1.19.709.883 I slot release: id 8 | task 411 | stop processing: n_tokens = 139, truncated = 0 1.19.714.917 I srv operator(): Chat format: peg-native 1.19.718.149 I srv operator(): Chat format: peg-native 1.19.797.468 I slot get_availabl: id 8 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.19.797.470 I srv get_availabl: updating prompt cache 1.19.797.513 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.19.797.515 I srv alloc: - prompt is already in the cache, skipping 1.19.797.516 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.19.797.519 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.19.797.521 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.19.797.521 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.19.797.522 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.19.797.522 I srv get_availabl: prompt cache update took 0.05 ms 1.19.797.589 I slot launch_slot_: id 8 | task 483 | processing task, is_child = 0 1.19.797.590 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 1.19.797.606 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.19.797.608 I srv alloc: - prompt is already in the cache, skipping 1.19.797.610 I slot get_availabl: id 10 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 1.19.797.611 I srv get_availabl: updating prompt cache 1.19.797.624 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.19.797.625 I srv alloc: - prompt is already in the cache, skipping 1.19.797.626 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 1.19.797.627 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.19.797.627 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.19.797.627 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.19.797.628 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.19.797.628 I srv get_availabl: prompt cache update took 0.02 ms 1.19.797.647 I slot launch_slot_: id 10 | task 484 | processing task, is_child = 0 1.19.800.607 I slot operator(): id 8 | task 483 | Checking checkpoint with [20, 20] against 24... 1.19.804.505 W slot operator(): id 8 | task 483 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.19.804.540 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.19.804.552 I slot operator(): id 10 | task 484 | Checking checkpoint with [23, 23] against 3... 1.19.804.554 W slot operator(): id 10 | task 484 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.19.804.555 W slot operator(): id 10 | task 484 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.19.806.006 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.19.908.443 I slot print_timing: id 12 | task 378 | prompt eval time = 94.09 ms / 4 tokens ( 23.52 ms per token, 42.51 tokens per second) 1.19.908.446 I slot print_timing: id 12 | task 378 | eval time = 8568.85 ms / 128 tokens ( 66.94 ms per token, 14.94 tokens per second) 1.19.908.447 I slot print_timing: id 12 | task 378 | total time = 8662.94 ms / 132 tokens 1.19.908.447 I slot print_timing: id 12 | task 378 | graphs reused = 1 1.19.908.450 I slot print_timing: id 12 | task 378 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.19.908.464 I statistics draft-mtp: #calls(b,g,a) = 91 392 6025, #gen drafts = 6038, #acc drafts = 4138, #gen tokens = 6038, #acc tokens = 4138, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.067, 667.165, 7.243 ms 1.19.908.491 I slot release: id 12 | task 378 | stop processing: n_tokens = 156, truncated = 0 1.19.916.983 I srv operator(): Chat format: peg-native 1.19.930.200 I slot create_check: id 10 | task 484 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.20.016.730 I slot print_timing: id 3 | task 418 | n_decoded = 100, tg = 18.69 t/s, tg_3s = 18.69 t/s 1.20.020.727 I slot get_availabl: id 12 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 1.20.020.728 I srv get_availabl: updating prompt cache 1.20.020.770 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.20.020.772 I srv alloc: - prompt is already in the cache, skipping 1.20.020.773 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 1.20.020.776 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.20.020.777 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.20.020.778 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.20.020.778 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.20.020.779 I srv get_availabl: prompt cache update took 0.05 ms 1.20.020.849 I slot launch_slot_: id 12 | task 487 | processing task, is_child = 0 1.20.022.853 I slot operator(): id 12 | task 487 | Checking checkpoint with [24, 24] against 3... 1.20.022.856 W slot operator(): id 12 | task 487 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.20.022.859 W slot operator(): id 12 | task 487 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.20.024.483 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.20.145.883 I slot create_check: id 12 | task 487 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.20.524.709 I slot print_timing: id 6 | task 428 | n_decoded = 101, tg = 20.29 t/s, tg_3s = 20.29 t/s 1.20.713.876 I slot print_timing: id 1 | task 412 | n_decoded = 101, tg = 15.20 t/s, tg_3s = 15.20 t/s 1.21.188.815 I slot print_timing: id 2 | task 410 | prompt eval time = 102.17 ms / 4 tokens ( 25.54 ms per token, 39.15 tokens per second) 1.21.188.818 I slot print_timing: id 2 | task 410 | eval time = 7120.52 ms / 128 tokens ( 55.63 ms per token, 17.98 tokens per second) 1.21.188.819 I slot print_timing: id 2 | task 410 | total time = 7222.69 ms / 132 tokens 1.21.188.820 I slot print_timing: id 2 | task 410 | graphs reused = 1 1.21.188.823 I slot print_timing: id 2 | task 410 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.21.188.836 I statistics draft-mtp: #calls(b,g,a) = 93 405 6226, #gen drafts = 6241, #acc drafts = 4275, #gen tokens = 6241, #acc tokens = 4275, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.069, 688.566, 7.492 ms 1.21.188.865 I slot release: id 2 | task 410 | stop processing: n_tokens = 155, truncated = 0 1.21.191.414 I slot print_timing: id 6 | task 428 | prompt eval time = 210.06 ms / 25 tokens ( 8.40 ms per token, 119.01 tokens per second) 1.21.191.416 I slot print_timing: id 6 | task 428 | eval time = 5644.62 ms / 114 tokens ( 49.51 ms per token, 20.20 tokens per second) 1.21.191.416 I slot print_timing: id 6 | task 428 | total time = 5854.68 ms / 139 tokens 1.21.191.417 I slot print_timing: id 6 | task 428 | graphs reused = 1 1.21.191.419 I slot print_timing: id 6 | task 428 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.21.191.425 I statistics draft-mtp: #calls(b,g,a) = 93 405 6232, #gen drafts = 6241, #acc drafts = 4280, #gen tokens = 6241, #acc tokens = 4280, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.069, 688.566, 7.499 ms 1.21.191.446 I slot release: id 6 | task 428 | stop processing: n_tokens = 139, truncated = 0 1.21.198.083 I srv operator(): Chat format: peg-native 1.21.201.303 I srv operator(): Chat format: peg-native 1.21.280.386 I slot get_availabl: id 6 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.21.280.388 I srv get_availabl: updating prompt cache 1.21.280.434 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.21.280.436 I srv alloc: - prompt is already in the cache, skipping 1.21.280.437 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.21.280.440 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.21.280.442 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.21.280.442 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.21.280.442 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.21.280.443 I srv get_availabl: prompt cache update took 0.05 ms 1.21.280.512 I slot launch_slot_: id 6 | task 501 | processing task, is_child = 0 1.21.280.513 I slot process_sing: id 2 | task -1 | saving idle slot to prompt cache 1.21.280.530 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.21.280.532 I srv alloc: - prompt is already in the cache, skipping 1.21.280.534 I slot get_availabl: id 2 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 1.21.280.534 I srv get_availabl: updating prompt cache 1.21.280.548 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.21.280.549 I srv alloc: - prompt is already in the cache, skipping 1.21.280.550 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 1.21.280.551 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.21.280.551 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.21.280.551 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.21.280.552 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.21.280.553 I srv get_availabl: prompt cache update took 0.02 ms 1.21.280.571 I slot launch_slot_: id 2 | task 502 | processing task, is_child = 0 1.21.283.001 I slot operator(): id 2 | task 502 | Checking checkpoint with [23, 23] against 3... 1.21.283.004 W slot operator(): id 2 | task 502 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.21.283.006 W slot operator(): id 2 | task 502 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.21.284.738 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.21.284.766 I slot operator(): id 6 | task 501 | Checking checkpoint with [20, 20] against 24... 1.21.288.720 W slot operator(): id 6 | task 501 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.21.288.755 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.21.413.709 I slot create_check: id 2 | task 502 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.21.598.580 I slot print_timing: id 9 | task 397 | prompt eval time = 215.19 ms / 29 tokens ( 7.42 ms per token, 134.76 tokens per second) 1.21.598.584 I slot print_timing: id 9 | task 397 | eval time = 8581.57 ms / 128 tokens ( 67.04 ms per token, 14.92 tokens per second) 1.21.598.584 I slot print_timing: id 9 | task 397 | total time = 8796.76 ms / 157 tokens 1.21.598.585 I slot print_timing: id 9 | task 397 | graphs reused = 1 1.21.598.587 I slot print_timing: id 9 | task 397 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.21.598.603 I statistics draft-mtp: #calls(b,g,a) = 95 409 6284, #gen drafts = 6299, #acc drafts = 4313, #gen tokens = 6299, #acc tokens = 4313, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.071, 697.459, 7.566 ms 1.21.598.635 I slot release: id 9 | task 397 | stop processing: n_tokens = 156, truncated = 0 1.21.607.220 I srv operator(): Chat format: peg-native 1.21.695.254 I slot get_availabl: id 9 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 1.21.695.256 I srv get_availabl: updating prompt cache 1.21.695.297 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.21.695.300 I srv alloc: - prompt is already in the cache, skipping 1.21.695.301 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 1.21.695.304 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.21.695.305 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.21.695.306 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.21.695.307 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.21.695.307 I srv get_availabl: prompt cache update took 0.05 ms 1.21.695.379 I slot launch_slot_: id 9 | task 507 | processing task, is_child = 0 1.21.697.999 I slot operator(): id 9 | task 507 | Checking checkpoint with [24, 24] against 3... 1.21.698.001 W slot operator(): id 9 | task 507 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.21.698.004 W slot operator(): id 9 | task 507 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.21.699.780 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.21.799.242 I slot print_timing: id 3 | task 418 | prompt eval time = 213.36 ms / 28 tokens ( 7.62 ms per token, 131.24 tokens per second) 1.21.799.245 I slot print_timing: id 3 | task 418 | eval time = 7133.59 ms / 128 tokens ( 55.73 ms per token, 17.94 tokens per second) 1.21.799.246 I slot print_timing: id 3 | task 418 | total time = 7346.95 ms / 156 tokens 1.21.799.246 I slot print_timing: id 3 | task 418 | graphs reused = 1 1.21.799.249 I slot print_timing: id 3 | task 418 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.21.799.268 I statistics draft-mtp: #calls(b,g,a) = 95 411 6314, #gen drafts = 6328, #acc drafts = 4333, #gen tokens = 6328, #acc tokens = 4333, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.071, 702.154, 7.606 ms 1.21.799.304 I slot release: id 3 | task 418 | stop processing: n_tokens = 155, truncated = 0 1.21.808.203 I srv operator(): Chat format: peg-native 1.21.821.770 I slot create_check: id 9 | task 507 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.21.913.071 I slot get_availabl: id 3 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.21.913.073 I srv get_availabl: updating prompt cache 1.21.913.114 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.21.913.116 I srv alloc: - prompt is already in the cache, skipping 1.21.913.117 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.21.913.120 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.21.913.121 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.21.913.121 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.21.913.122 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.21.913.122 I srv get_availabl: prompt cache update took 0.05 ms 1.21.913.191 I slot launch_slot_: id 3 | task 510 | processing task, is_child = 0 1.21.915.217 I slot operator(): id 3 | task 510 | Checking checkpoint with [23, 23] against 3... 1.21.915.220 W slot operator(): id 3 | task 510 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.21.915.221 W slot operator(): id 3 | task 510 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.21.917.100 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.22.038.289 I slot create_check: id 3 | task 510 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.22.131.366 I slot print_timing: id 7 | task 448 | n_decoded = 101, tg = 20.09 t/s, tg_3s = 20.09 t/s 1.22.419.352 I slot print_timing: id 13 | task 431 | n_decoded = 101, tg = 15.19 t/s, tg_3s = 15.19 t/s 1.22.419.828 I slot print_timing: id 14 | task 446 | n_decoded = 100, tg = 18.44 t/s, tg_3s = 18.44 t/s 1.22.699.165 I slot print_timing: id 1 | task 412 | prompt eval time = 105.87 ms / 4 tokens ( 26.47 ms per token, 37.78 tokens per second) 1.22.699.168 I slot print_timing: id 1 | task 412 | eval time = 8631.13 ms / 128 tokens ( 67.43 ms per token, 14.83 tokens per second) 1.22.699.168 I slot print_timing: id 1 | task 412 | total time = 8737.00 ms / 132 tokens 1.22.699.169 I slot print_timing: id 1 | task 412 | graphs reused = 1 1.22.699.172 I slot print_timing: id 1 | task 412 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.22.699.188 I statistics draft-mtp: #calls(b,g,a) = 97 420 6452, #gen drafts = 6467, #acc drafts = 4429, #gen tokens = 6467, #acc tokens = 4429, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.072, 718.000, 7.770 ms 1.22.699.217 I slot release: id 1 | task 412 | stop processing: n_tokens = 156, truncated = 0 1.22.708.309 I srv operator(): Chat format: peg-native 1.22.792.543 I slot print_timing: id 7 | task 448 | prompt eval time = 99.56 ms / 4 tokens ( 24.89 ms per token, 40.18 tokens per second) 1.22.792.546 I slot print_timing: id 7 | task 448 | eval time = 5689.26 ms / 114 tokens ( 49.91 ms per token, 20.04 tokens per second) 1.22.792.546 I slot print_timing: id 7 | task 448 | total time = 5788.82 ms / 118 tokens 1.22.792.547 I slot print_timing: id 7 | task 448 | graphs reused = 1 1.22.792.550 I slot print_timing: id 7 | task 448 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.22.792.562 I statistics draft-mtp: #calls(b,g,a) = 97 421 6474, #gen drafts = 6482, #acc drafts = 4445, #gen tokens = 6482, #acc tokens = 4445, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.072, 719.870, 7.797 ms 1.22.792.593 I slot release: id 7 | task 448 | stop processing: n_tokens = 139, truncated = 0 1.22.796.003 I slot get_availabl: id 1 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.22.796.005 I srv get_availabl: updating prompt cache 1.22.796.046 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.22.796.049 I srv alloc: - prompt is already in the cache, skipping 1.22.796.049 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.22.796.053 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.22.796.053 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.22.796.054 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.22.796.055 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.22.796.055 I srv get_availabl: prompt cache update took 0.05 ms 1.22.796.125 I slot launch_slot_: id 1 | task 520 | processing task, is_child = 0 1.22.796.127 I slot process_sing: id 7 | task -1 | saving idle slot to prompt cache 1.22.796.143 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.22.796.144 I srv alloc: - prompt is already in the cache, skipping 1.22.798.851 I slot operator(): id 1 | task 520 | Checking checkpoint with [24, 24] against 28... 1.22.801.639 I srv operator(): Chat format: peg-native 1.22.802.788 W slot operator(): id 1 | task 520 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.22.802.817 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.22.893.528 I slot get_availabl: id 7 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.022 1.22.893.531 I srv get_availabl: updating prompt cache 1.22.893.572 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.22.893.574 I srv alloc: - prompt is already in the cache, skipping 1.22.893.575 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.107 1.22.893.578 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.22.893.579 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.22.893.580 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.22.893.580 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.22.893.581 I srv get_availabl: prompt cache update took 0.05 ms 1.22.893.647 I slot launch_slot_: id 7 | task 522 | processing task, is_child = 0 1.22.896.061 I slot operator(): id 7 | task 522 | Checking checkpoint with [20, 20] against 3... 1.22.896.064 W slot operator(): id 7 | task 522 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.22.896.065 W slot operator(): id 7 | task 522 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.22.897.622 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.23.020.148 I slot create_check: id 7 | task 522 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.23.588.653 I slot print_timing: id 5 | task 466 | n_decoded = 101, tg = 20.02 t/s, tg_3s = 20.02 t/s 1.23.876.873 I slot print_timing: id 11 | task 464 | n_decoded = 100, tg = 18.34 t/s, tg_3s = 18.34 t/s 1.23.878.345 I slot print_timing: id 15 | task 449 | n_decoded = 101, tg = 14.91 t/s, tg_3s = 14.91 t/s 1.24.065.112 I slot print_timing: id 4 | task 468 | n_decoded = 100, tg = 18.49 t/s, tg_3s = 18.49 t/s 1.24.158.751 I slot print_timing: id 14 | task 446 | prompt eval time = 83.83 ms / 4 tokens ( 20.96 ms per token, 47.72 tokens per second) 1.24.158.754 I slot print_timing: id 14 | task 446 | eval time = 7162.24 ms / 128 tokens ( 55.95 ms per token, 17.87 tokens per second) 1.24.158.754 I slot print_timing: id 14 | task 446 | total time = 7246.07 ms / 132 tokens 1.24.158.755 I slot print_timing: id 14 | task 446 | graphs reused = 1 1.24.158.758 I slot print_timing: id 14 | task 446 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.24.158.774 I statistics draft-mtp: #calls(b,g,a) = 99 435 6686, #gen drafts = 6701, #acc drafts = 4592, #gen tokens = 6701, #acc tokens = 4592, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.074, 743.789, 8.047 ms 1.24.158.802 I slot release: id 14 | task 446 | stop processing: n_tokens = 155, truncated = 0 1.24.168.030 I srv operator(): Chat format: peg-native 1.24.252.008 I slot print_timing: id 5 | task 466 | prompt eval time = 108.51 ms / 4 tokens ( 27.13 ms per token, 36.86 tokens per second) 1.24.252.011 I slot print_timing: id 5 | task 466 | eval time = 5707.58 ms / 114 tokens ( 50.07 ms per token, 19.97 tokens per second) 1.24.252.012 I slot print_timing: id 5 | task 466 | total time = 5816.09 ms / 118 tokens 1.24.252.013 I slot print_timing: id 5 | task 466 | graphs reused = 1 1.24.252.016 I slot print_timing: id 5 | task 466 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.24.252.027 I statistics draft-mtp: #calls(b,g,a) = 99 436 6707, #gen drafts = 6716, #acc drafts = 4610, #gen tokens = 6716, #acc tokens = 4610, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.074, 745.657, 8.074 ms 1.24.252.053 I slot release: id 5 | task 466 | stop processing: n_tokens = 139, truncated = 0 1.24.255.429 I slot get_availabl: id 5 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.24.255.431 I srv get_availabl: updating prompt cache 1.24.255.476 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.24.255.478 I srv alloc: - prompt is already in the cache, skipping 1.24.255.479 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.24.255.483 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.24.255.485 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.24.255.485 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.24.255.486 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.24.255.486 I srv get_availabl: prompt cache update took 0.06 ms 1.24.255.558 I slot launch_slot_: id 5 | task 537 | processing task, is_child = 0 1.24.255.561 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 1.24.255.578 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.24.255.578 I srv alloc: - prompt is already in the cache, skipping 1.24.258.130 I slot operator(): id 5 | task 537 | Checking checkpoint with [20, 20] against 24... 1.24.260.271 I srv operator(): Chat format: peg-native 1.24.262.049 W slot operator(): id 5 | task 537 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.24.262.087 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.24.346.519 I slot print_timing: id 13 | task 431 | prompt eval time = 214.15 ms / 29 tokens ( 7.38 ms per token, 135.42 tokens per second) 1.24.346.522 I slot print_timing: id 13 | task 431 | eval time = 8577.95 ms / 128 tokens ( 67.02 ms per token, 14.92 tokens per second) 1.24.346.523 I slot print_timing: id 13 | task 431 | total time = 8792.10 ms / 157 tokens 1.24.346.523 I slot print_timing: id 13 | task 431 | graphs reused = 1 1.24.346.526 I slot print_timing: id 13 | task 431 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.24.346.540 I statistics draft-mtp: #calls(b,g,a) = 100 437 6716, #gen drafts = 6729, #acc drafts = 4615, #gen tokens = 6729, #acc tokens = 4615, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.074, 748.173, 8.084 ms 1.24.346.569 I slot release: id 13 | task 431 | stop processing: n_tokens = 156, truncated = 0 1.24.351.815 I slot get_availabl: id 13 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.24.351.817 I srv get_availabl: updating prompt cache 1.24.351.861 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.24.351.864 I srv alloc: - prompt is already in the cache, skipping 1.24.351.865 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.24.351.867 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.24.351.868 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.24.351.869 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.24.351.869 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.24.351.870 I srv get_availabl: prompt cache update took 0.05 ms 1.24.351.938 I slot launch_slot_: id 13 | task 539 | processing task, is_child = 0 1.24.351.940 I slot process_sing: id 14 | task -1 | saving idle slot to prompt cache 1.24.351.956 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.24.351.957 I srv alloc: - prompt is already in the cache, skipping 1.24.353.789 I slot operator(): id 13 | task 539 | Checking checkpoint with [24, 24] against 28... 1.24.355.026 I srv operator(): Chat format: peg-native 1.24.357.750 W slot operator(): id 13 | task 539 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.24.357.784 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.24.448.373 I slot get_availabl: id 14 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.24.448.376 I srv get_availabl: updating prompt cache 1.24.448.420 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.24.448.422 I srv alloc: - prompt is already in the cache, skipping 1.24.448.423 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.24.448.426 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.24.448.427 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.24.448.428 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.24.448.428 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.24.448.429 I srv get_availabl: prompt cache update took 0.05 ms 1.24.448.496 I slot launch_slot_: id 14 | task 541 | processing task, is_child = 0 1.24.450.410 I slot operator(): id 14 | task 541 | Checking checkpoint with [23, 23] against 27... 1.24.454.354 W slot operator(): id 14 | task 541 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.24.454.373 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.24.929.381 I slot print_timing: id 8 | task 483 | n_decoded = 101, tg = 20.11 t/s, tg_3s = 20.11 t/s 1.25.212.464 I slot print_timing: id 0 | task 467 | n_decoded = 101, tg = 15.15 t/s, tg_3s = 15.15 t/s 1.25.594.762 I slot print_timing: id 11 | task 464 | prompt eval time = 77.90 ms / 4 tokens ( 19.48 ms per token, 51.35 tokens per second) 1.25.594.766 I slot print_timing: id 11 | task 464 | eval time = 7171.34 ms / 128 tokens ( 56.03 ms per token, 17.85 tokens per second) 1.25.594.766 I slot print_timing: id 11 | task 464 | total time = 7249.24 ms / 132 tokens 1.25.594.767 I slot print_timing: id 11 | task 464 | graphs reused = 1 1.25.594.770 I slot print_timing: id 11 | task 464 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.25.594.785 I statistics draft-mtp: #calls(b,g,a) = 102 450 6918, #gen drafts = 6933, #acc drafts = 4751, #gen tokens = 6933, #acc tokens = 4751, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.076, 768.660, 8.335 ms 1.25.594.813 I slot release: id 11 | task 464 | stop processing: n_tokens = 155, truncated = 0 1.25.598.710 I slot print_timing: id 8 | task 483 | prompt eval time = 107.58 ms / 4 tokens ( 26.89 ms per token, 37.18 tokens per second) 1.25.598.713 I slot print_timing: id 8 | task 483 | eval time = 5690.51 ms / 114 tokens ( 49.92 ms per token, 20.03 tokens per second) 1.25.598.713 I slot print_timing: id 8 | task 483 | total time = 5798.09 ms / 118 tokens 1.25.598.714 I slot print_timing: id 8 | task 483 | graphs reused = 1 1.25.598.716 I slot print_timing: id 8 | task 483 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.25.598.723 I statistics draft-mtp: #calls(b,g,a) = 102 450 6927, #gen drafts = 6933, #acc drafts = 4759, #gen tokens = 6933, #acc tokens = 4759, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.076, 768.660, 8.344 ms 1.25.598.746 I slot release: id 8 | task 483 | stop processing: n_tokens = 139, truncated = 0 1.25.599.767 I slot print_timing: id 12 | task 487 | n_decoded = 100, tg = 18.64 t/s, tg_3s = 18.64 t/s 1.25.603.870 I srv operator(): Chat format: peg-native 1.25.607.210 I srv operator(): Chat format: peg-native 1.25.686.547 I slot get_availabl: id 8 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.25.686.550 I srv get_availabl: updating prompt cache 1.25.686.592 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.25.686.595 I srv alloc: - prompt is already in the cache, skipping 1.25.686.596 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.25.686.599 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.25.686.599 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.25.686.600 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.25.686.601 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.25.686.602 I srv get_availabl: prompt cache update took 0.05 ms 1.25.686.671 I slot launch_slot_: id 8 | task 555 | processing task, is_child = 0 1.25.686.676 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 1.25.686.693 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.25.686.694 I srv alloc: - prompt is already in the cache, skipping 1.25.686.696 I slot get_availabl: id 11 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 1.25.686.696 I srv get_availabl: updating prompt cache 1.25.686.710 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.25.686.711 I srv alloc: - prompt is already in the cache, skipping 1.25.686.712 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 1.25.686.712 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.25.686.713 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.25.686.713 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.25.686.713 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.25.686.713 I srv get_availabl: prompt cache update took 0.02 ms 1.25.686.730 I slot launch_slot_: id 11 | task 556 | processing task, is_child = 0 1.25.689.581 I slot operator(): id 8 | task 555 | Checking checkpoint with [20, 20] against 24... 1.25.693.462 W slot operator(): id 8 | task 555 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.25.693.497 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.25.693.517 I slot operator(): id 11 | task 556 | Checking checkpoint with [23, 23] against 3... 1.25.693.519 W slot operator(): id 11 | task 556 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.25.693.520 W slot operator(): id 11 | task 556 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.25.695.129 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.25.796.234 I slot print_timing: id 4 | task 468 | prompt eval time = 223.17 ms / 28 tokens ( 7.97 ms per token, 125.46 tokens per second) 1.25.796.237 I slot print_timing: id 4 | task 468 | eval time = 7138.88 ms / 128 tokens ( 55.77 ms per token, 17.93 tokens per second) 1.25.796.238 I slot print_timing: id 4 | task 468 | total time = 7362.05 ms / 156 tokens 1.25.796.239 I slot print_timing: id 4 | task 468 | graphs reused = 1 1.25.796.241 I slot print_timing: id 4 | task 468 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.25.796.257 I statistics draft-mtp: #calls(b,g,a) = 102 452 6947, #gen drafts = 6959, #acc drafts = 4771, #gen tokens = 6959, #acc tokens = 4771, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.076, 774.152, 8.367 ms 1.25.796.285 I slot release: id 4 | task 468 | stop processing: n_tokens = 155, truncated = 0 1.25.796.807 I slot print_timing: id 15 | task 449 | prompt eval time = 95.94 ms / 4 tokens ( 23.99 ms per token, 41.69 tokens per second) 1.25.796.809 I slot print_timing: id 15 | task 449 | eval time = 8693.24 ms / 128 tokens ( 67.92 ms per token, 14.72 tokens per second) 1.25.796.809 I slot print_timing: id 15 | task 449 | total time = 8789.19 ms / 132 tokens 1.25.796.810 I slot print_timing: id 15 | task 449 | graphs reused = 1 1.25.796.811 I slot print_timing: id 15 | task 449 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.25.796.815 I statistics draft-mtp: #calls(b,g,a) = 103 452 6947, #gen drafts = 6959, #acc drafts = 4771, #gen tokens = 6959, #acc tokens = 4771, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.076, 774.152, 8.367 ms 1.25.796.837 I slot release: id 15 | task 449 | stop processing: n_tokens = 156, truncated = 0 1.25.805.519 I srv operator(): Chat format: peg-native 1.25.805.874 I srv operator(): Chat format: peg-native 1.25.818.566 I slot create_check: id 11 | task 556 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.25.903.097 I slot get_availabl: id 4 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.25.903.099 I srv get_availabl: updating prompt cache 1.25.903.140 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.25.903.143 I srv alloc: - prompt is already in the cache, skipping 1.25.903.144 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.25.903.146 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.25.903.147 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.25.903.148 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.25.903.149 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.25.903.150 I srv get_availabl: prompt cache update took 0.05 ms 1.25.903.220 I slot launch_slot_: id 4 | task 559 | processing task, is_child = 0 1.25.903.221 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 1.25.903.239 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.25.903.240 I srv alloc: - prompt is already in the cache, skipping 1.25.903.242 I slot get_availabl: id 15 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.25.903.242 I srv get_availabl: updating prompt cache 1.25.903.256 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.25.903.257 I srv alloc: - prompt is already in the cache, skipping 1.25.903.258 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.25.903.258 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.25.903.258 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.25.903.259 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.25.903.259 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.25.903.259 I srv get_availabl: prompt cache update took 0.02 ms 1.25.903.278 I slot launch_slot_: id 15 | task 560 | processing task, is_child = 0 1.25.905.446 I slot operator(): id 4 | task 559 | Checking checkpoint with [23, 23] against 27... 1.25.909.320 W slot operator(): id 4 | task 559 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.25.909.356 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.25.909.370 I slot operator(): id 15 | task 560 | Checking checkpoint with [24, 24] against 3... 1.25.909.371 W slot operator(): id 15 | task 560 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.25.909.372 W slot operator(): id 15 | task 560 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.25.911.317 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.26.033.784 I slot create_check: id 15 | task 560 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.26.413.587 I slot print_timing: id 6 | task 501 | n_decoded = 101, tg = 20.11 t/s, tg_3s = 20.11 t/s 1.26.796.827 I slot print_timing: id 10 | task 484 | n_decoded = 101, tg = 14.89 t/s, tg_3s = 14.89 t/s 1.27.080.187 I slot print_timing: id 3 | task 510 | n_decoded = 101, tg = 20.40 t/s, tg_3s = 20.40 t/s 1.27.081.514 I slot print_timing: id 6 | task 501 | prompt eval time = 107.39 ms / 4 tokens ( 26.85 ms per token, 37.25 tokens per second) 1.27.081.516 I slot print_timing: id 6 | task 501 | eval time = 5689.35 ms / 114 tokens ( 49.91 ms per token, 20.04 tokens per second) 1.27.081.516 I slot print_timing: id 6 | task 501 | total time = 5796.73 ms / 118 tokens 1.27.081.517 I slot print_timing: id 6 | task 501 | graphs reused = 1 1.27.081.519 I slot print_timing: id 6 | task 501 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.27.081.535 I statistics draft-mtp: #calls(b,g,a) = 106 465 7152, #gen drafts = 7161, #acc drafts = 4910, #gen tokens = 7161, #acc tokens = 4910, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.079, 794.960, 8.603 ms 1.27.081.565 I slot release: id 6 | task 501 | stop processing: n_tokens = 139, truncated = 0 1.27.089.989 I srv operator(): Chat format: peg-native 1.27.168.228 I slot print_timing: id 0 | task 467 | prompt eval time = 113.91 ms / 4 tokens ( 28.48 ms per token, 35.12 tokens per second) 1.27.168.231 I slot print_timing: id 0 | task 467 | eval time = 8624.07 ms / 128 tokens ( 67.38 ms per token, 14.84 tokens per second) 1.27.168.232 I slot print_timing: id 0 | task 467 | total time = 8737.98 ms / 132 tokens 1.27.168.233 I slot print_timing: id 0 | task 467 | graphs reused = 1 1.27.168.235 I slot print_timing: id 0 | task 467 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.27.168.249 I statistics draft-mtp: #calls(b,g,a) = 106 466 7161, #gen drafts = 7175, #acc drafts = 4915, #gen tokens = 7175, #acc tokens = 4915, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.079, 796.961, 8.613 ms 1.27.168.275 I slot release: id 0 | task 467 | stop processing: n_tokens = 156, truncated = 0 1.27.174.315 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.27.174.317 I srv get_availabl: updating prompt cache 1.27.174.360 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.27.174.362 I srv alloc: - prompt is already in the cache, skipping 1.27.174.363 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.27.174.366 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.27.174.367 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.27.174.367 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.27.174.368 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.27.174.368 I srv get_availabl: prompt cache update took 0.05 ms 1.27.174.445 I slot launch_slot_: id 0 | task 574 | processing task, is_child = 0 1.27.174.449 I slot process_sing: id 6 | task -1 | saving idle slot to prompt cache 1.27.174.466 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.27.174.467 I srv alloc: - prompt is already in the cache, skipping 1.27.176.504 I slot operator(): id 0 | task 574 | Checking checkpoint with [24, 24] against 28... 1.27.177.902 I srv operator(): Chat format: peg-native 1.27.180.414 W slot operator(): id 0 | task 574 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.27.180.466 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.27.268.535 I slot print_timing: id 9 | task 507 | n_decoded = 100, tg = 18.65 t/s, tg_3s = 18.65 t/s 1.27.270.743 I slot get_availabl: id 6 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.022 1.27.270.745 I srv get_availabl: updating prompt cache 1.27.270.786 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.27.270.790 I srv alloc: - prompt is already in the cache, skipping 1.27.270.791 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.107 1.27.270.794 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.27.270.796 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.27.270.797 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.27.270.797 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.27.270.798 I srv get_availabl: prompt cache update took 0.05 ms 1.27.270.869 I slot launch_slot_: id 6 | task 576 | processing task, is_child = 0 1.27.273.439 I slot operator(): id 6 | task 576 | Checking checkpoint with [20, 20] against 3... 1.27.273.442 W slot operator(): id 6 | task 576 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.27.273.445 W slot operator(): id 6 | task 576 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.27.276.062 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.27.375.680 I slot print_timing: id 12 | task 487 | prompt eval time = 213.01 ms / 28 tokens ( 7.61 ms per token, 131.45 tokens per second) 1.27.375.683 I slot print_timing: id 12 | task 487 | eval time = 7139.79 ms / 128 tokens ( 55.78 ms per token, 17.93 tokens per second) 1.27.375.683 I slot print_timing: id 12 | task 487 | total time = 7352.80 ms / 156 tokens 1.27.375.684 I slot print_timing: id 12 | task 487 | graphs reused = 1 1.27.375.687 I slot print_timing: id 12 | task 487 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.27.375.702 I statistics draft-mtp: #calls(b,g,a) = 107 468 7189, #gen drafts = 7203, #acc drafts = 4935, #gen tokens = 7203, #acc tokens = 4935, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.080, 801.501, 8.649 ms 1.27.375.732 I slot release: id 12 | task 487 | stop processing: n_tokens = 155, truncated = 0 1.27.384.709 I srv operator(): Chat format: peg-native 1.27.398.470 I slot create_check: id 6 | task 576 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.27.488.652 I slot get_availabl: id 12 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.27.488.655 I srv get_availabl: updating prompt cache 1.27.488.695 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.27.488.698 I srv alloc: - prompt is already in the cache, skipping 1.27.488.699 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.27.488.702 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.27.488.703 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.27.488.704 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.27.488.705 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.27.488.705 I srv get_availabl: prompt cache update took 0.05 ms 1.27.488.776 I slot launch_slot_: id 12 | task 579 | processing task, is_child = 0 1.27.490.790 I slot operator(): id 12 | task 579 | Checking checkpoint with [23, 23] against 3... 1.27.490.793 W slot operator(): id 12 | task 579 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.27.490.795 W slot operator(): id 12 | task 579 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.27.492.429 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.27.612.954 I slot create_check: id 12 | task 579 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.27.799.963 I slot print_timing: id 3 | task 510 | prompt eval time = 213.05 ms / 25 tokens ( 8.52 ms per token, 117.34 tokens per second) 1.27.799.966 I slot print_timing: id 3 | task 510 | eval time = 5671.67 ms / 114 tokens ( 49.75 ms per token, 20.10 tokens per second) 1.27.799.966 I slot print_timing: id 3 | task 510 | total time = 5884.72 ms / 139 tokens 1.27.799.968 I slot print_timing: id 3 | task 510 | graphs reused = 1 1.27.799.970 I slot print_timing: id 3 | task 510 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.27.799.990 I statistics draft-mtp: #calls(b,g,a) = 109 472 7251, #gen drafts = 7263, #acc drafts = 4979, #gen tokens = 7263, #acc tokens = 4979, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.082, 809.461, 8.726 ms 1.27.800.023 I slot release: id 3 | task 510 | stop processing: n_tokens = 139, truncated = 0 1.27.808.864 I srv operator(): Chat format: peg-native 1.27.895.684 I slot get_availabl: id 3 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 1.27.895.687 I srv get_availabl: updating prompt cache 1.27.895.727 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.27.895.730 I srv alloc: - prompt is already in the cache, skipping 1.27.895.731 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 1.27.895.734 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.27.895.735 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.27.895.735 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.27.895.736 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.27.895.736 I srv get_availabl: prompt cache update took 0.05 ms 1.27.895.806 I slot launch_slot_: id 3 | task 584 | processing task, is_child = 0 1.27.897.826 I slot operator(): id 3 | task 584 | Checking checkpoint with [20, 20] against 3... 1.27.897.828 W slot operator(): id 3 | task 584 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.27.897.832 W slot operator(): id 3 | task 584 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.27.899.837 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.28.022.412 I slot create_check: id 3 | task 584 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.28.208.362 I slot print_timing: id 2 | task 502 | n_decoded = 101, tg = 15.06 t/s, tg_3s = 15.06 t/s 1.28.496.695 I slot print_timing: id 7 | task 522 | n_decoded = 100, tg = 18.57 t/s, tg_3s = 18.57 t/s 1.28.780.479 I slot print_timing: id 10 | task 484 | prompt eval time = 210.40 ms / 29 tokens ( 7.26 ms per token, 137.83 tokens per second) 1.28.780.482 I slot print_timing: id 10 | task 484 | eval time = 8765.50 ms / 128 tokens ( 68.48 ms per token, 14.60 tokens per second) 1.28.780.482 I slot print_timing: id 10 | task 484 | total time = 8975.90 ms / 157 tokens 1.28.780.484 I slot print_timing: id 10 | task 484 | graphs reused = 1 1.28.780.486 I slot print_timing: id 10 | task 484 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.28.780.501 I statistics draft-mtp: #calls(b,g,a) = 110 482 7404, #gen drafts = 7419, #acc drafts = 5085, #gen tokens = 7419, #acc tokens = 5085, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.083, 826.361, 8.911 ms 1.28.780.529 I slot release: id 10 | task 484 | stop processing: n_tokens = 156, truncated = 0 1.28.789.743 I srv operator(): Chat format: peg-native 1.28.877.438 I slot get_availabl: id 10 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 1.28.877.440 I srv get_availabl: updating prompt cache 1.28.877.484 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.28.877.486 I srv alloc: - prompt is already in the cache, skipping 1.28.877.488 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 1.28.877.492 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.28.877.493 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.28.877.494 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.28.877.494 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.28.877.494 I srv get_availabl: prompt cache update took 0.05 ms 1.28.877.562 I slot launch_slot_: id 10 | task 595 | processing task, is_child = 0 1.28.879.979 I slot operator(): id 10 | task 595 | Checking checkpoint with [24, 24] against 3... 1.28.879.981 W slot operator(): id 10 | task 595 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.28.879.983 W slot operator(): id 10 | task 595 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.28.881.548 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.29.003.208 I slot create_check: id 10 | task 595 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.29.092.593 I slot print_timing: id 9 | task 507 | prompt eval time = 209.13 ms / 28 tokens ( 7.47 ms per token, 133.89 tokens per second) 1.29.092.596 I slot print_timing: id 9 | task 507 | eval time = 7185.44 ms / 128 tokens ( 56.14 ms per token, 17.81 tokens per second) 1.29.092.596 I slot print_timing: id 9 | task 507 | total time = 7394.56 ms / 156 tokens 1.29.092.597 I slot print_timing: id 9 | task 507 | graphs reused = 1 1.29.092.600 I slot print_timing: id 9 | task 507 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.29.092.615 I statistics draft-mtp: #calls(b,g,a) = 110 485 7449, #gen drafts = 7463, #acc drafts = 5115, #gen tokens = 7463, #acc tokens = 5115, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.083, 832.743, 8.962 ms 1.29.092.643 I slot release: id 9 | task 507 | stop processing: n_tokens = 155, truncated = 0 1.29.101.728 I srv operator(): Chat format: peg-native 1.29.188.859 I slot get_availabl: id 9 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.29.188.862 I srv get_availabl: updating prompt cache 1.29.188.904 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.29.188.906 I srv alloc: - prompt is already in the cache, skipping 1.29.188.908 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.29.188.911 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.29.188.912 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.29.188.913 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.29.188.914 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.29.188.914 I srv get_availabl: prompt cache update took 0.05 ms 1.29.188.983 I slot launch_slot_: id 9 | task 599 | processing task, is_child = 0 1.29.191.494 I slot operator(): id 9 | task 599 | Checking checkpoint with [23, 23] against 3... 1.29.191.496 W slot operator(): id 9 | task 599 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.29.191.498 W slot operator(): id 9 | task 599 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.29.193.303 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.29.315.318 I slot create_check: id 9 | task 599 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.29.407.722 I slot print_timing: id 5 | task 537 | n_decoded = 101, tg = 19.95 t/s, tg_3s = 19.95 t/s 1.29.691.654 I slot print_timing: id 1 | task 520 | n_decoded = 101, tg = 14.84 t/s, tg_3s = 14.84 t/s 1.29.982.416 I slot print_timing: id 14 | task 541 | n_decoded = 100, tg = 18.39 t/s, tg_3s = 18.39 t/s 1.30.074.417 I slot print_timing: id 5 | task 537 | prompt eval time = 88.14 ms / 4 tokens ( 22.04 ms per token, 45.38 tokens per second) 1.30.074.421 I slot print_timing: id 5 | task 537 | eval time = 5728.12 ms / 114 tokens ( 50.25 ms per token, 19.90 tokens per second) 1.30.074.421 I slot print_timing: id 5 | task 537 | total time = 5816.26 ms / 118 tokens 1.30.074.422 I slot print_timing: id 5 | task 537 | graphs reused = 1 1.30.074.425 I slot print_timing: id 5 | task 537 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.30.074.443 I statistics draft-mtp: #calls(b,g,a) = 112 495 7610, #gen drafts = 7620, #acc drafts = 5225, #gen tokens = 7620, #acc tokens = 5225, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.085, 849.689, 9.162 ms 1.30.074.472 I slot release: id 5 | task 537 | stop processing: n_tokens = 139, truncated = 0 1.30.082.924 I srv operator(): Chat format: peg-native 1.30.162.484 I slot print_timing: id 2 | task 502 | prompt eval time = 220.59 ms / 29 tokens ( 7.61 ms per token, 131.47 tokens per second) 1.30.162.486 I slot print_timing: id 2 | task 502 | eval time = 8658.87 ms / 128 tokens ( 67.65 ms per token, 14.78 tokens per second) 1.30.162.487 I slot print_timing: id 2 | task 502 | total time = 8879.46 ms / 157 tokens 1.30.162.488 I slot print_timing: id 2 | task 502 | graphs reused = 1 1.30.162.490 I slot print_timing: id 2 | task 502 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.30.162.504 I statistics draft-mtp: #calls(b,g,a) = 112 496 7620, #gen drafts = 7634, #acc drafts = 5233, #gen tokens = 7634, #acc tokens = 5233, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.085, 852.011, 9.173 ms 1.30.162.533 I slot release: id 2 | task 502 | stop processing: n_tokens = 156, truncated = 0 1.30.168.058 I slot get_availabl: id 2 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.30.168.060 I srv get_availabl: updating prompt cache 1.30.168.104 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.30.168.107 I srv alloc: - prompt is already in the cache, skipping 1.30.168.108 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.30.168.112 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.30.168.113 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.30.168.114 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.30.168.114 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.30.168.114 I srv get_availabl: prompt cache update took 0.05 ms 1.30.168.186 I slot launch_slot_: id 2 | task 610 | processing task, is_child = 0 1.30.168.187 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 1.30.168.202 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.30.168.203 I srv alloc: - prompt is already in the cache, skipping 1.30.171.135 I slot operator(): id 2 | task 610 | Checking checkpoint with [24, 24] against 28... 1.30.171.287 I srv operator(): Chat format: peg-native 1.30.175.023 W slot operator(): id 2 | task 610 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.30.175.061 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.30.259.515 I slot print_timing: id 7 | task 522 | prompt eval time = 214.20 ms / 28 tokens ( 7.65 ms per token, 130.72 tokens per second) 1.30.259.517 I slot print_timing: id 7 | task 522 | eval time = 7149.23 ms / 128 tokens ( 55.85 ms per token, 17.90 tokens per second) 1.30.259.517 I slot print_timing: id 7 | task 522 | total time = 7363.44 ms / 156 tokens 1.30.259.518 I slot print_timing: id 7 | task 522 | graphs reused = 1 1.30.259.521 I slot print_timing: id 7 | task 522 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.30.259.535 I statistics draft-mtp: #calls(b,g,a) = 113 497 7634, #gen drafts = 7647, #acc drafts = 5242, #gen tokens = 7647, #acc tokens = 5242, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.086, 854.904, 9.192 ms 1.30.259.563 I slot release: id 7 | task 522 | stop processing: n_tokens = 155, truncated = 0 1.30.264.867 I slot get_availabl: id 7 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.30.264.868 I srv get_availabl: updating prompt cache 1.30.264.912 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.30.264.914 I srv alloc: - prompt is already in the cache, skipping 1.30.264.915 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.30.264.918 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.30.264.919 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.30.264.919 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.30.264.919 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.30.264.920 I srv get_availabl: prompt cache update took 0.05 ms 1.30.264.994 I slot launch_slot_: id 7 | task 612 | processing task, is_child = 0 1.30.264.995 I slot process_sing: id 5 | task -1 | saving idle slot to prompt cache 1.30.265.012 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.30.265.013 I srv alloc: - prompt is already in the cache, skipping 1.30.267.699 I slot operator(): id 7 | task 612 | Checking checkpoint with [23, 23] against 27... 1.30.268.556 I srv operator(): Chat format: peg-native 1.30.271.590 W slot operator(): id 7 | task 612 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.30.271.619 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.30.362.100 I slot get_availabl: id 5 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.30.362.103 I srv get_availabl: updating prompt cache 1.30.362.146 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.30.362.148 I srv alloc: - prompt is already in the cache, skipping 1.30.362.149 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.30.362.152 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.30.362.153 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.30.362.154 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.30.362.154 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.30.362.155 I srv get_availabl: prompt cache update took 0.05 ms 1.30.362.222 I slot launch_slot_: id 5 | task 614 | processing task, is_child = 0 1.30.364.342 I slot operator(): id 5 | task 614 | Checking checkpoint with [20, 20] against 24... 1.30.368.306 W slot operator(): id 5 | task 614 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.30.368.331 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.30.847.592 I slot print_timing: id 8 | task 555 | n_decoded = 101, tg = 20.00 t/s, tg_3s = 20.00 t/s 1.31.136.626 I slot print_timing: id 15 | task 560 | n_decoded = 101, tg = 20.15 t/s, tg_3s = 20.15 t/s 1.31.231.570 I slot print_timing: id 13 | task 539 | n_decoded = 101, tg = 14.88 t/s, tg_3s = 14.88 t/s 1.31.418.423 I slot print_timing: id 4 | task 559 | n_decoded = 100, tg = 18.50 t/s, tg_3s = 18.50 t/s 1.31.515.471 I slot print_timing: id 8 | task 555 | prompt eval time = 107.01 ms / 4 tokens ( 26.75 ms per token, 37.38 tokens per second) 1.31.515.474 I slot print_timing: id 8 | task 555 | eval time = 5718.85 ms / 114 tokens ( 50.17 ms per token, 19.93 tokens per second) 1.31.515.475 I slot print_timing: id 8 | task 555 | total time = 5825.87 ms / 118 tokens 1.31.515.476 I slot print_timing: id 8 | task 555 | graphs reused = 1 1.31.515.478 I slot print_timing: id 8 | task 555 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.31.515.497 I statistics draft-mtp: #calls(b,g,a) = 115 510 7845, #gen drafts = 7852, #acc drafts = 5391, #gen tokens = 7852, #acc tokens = 5391, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.088, 875.897, 9.442 ms 1.31.515.527 I slot release: id 8 | task 555 | stop processing: n_tokens = 139, truncated = 0 1.31.524.116 I srv operator(): Chat format: peg-native 1.31.602.873 I slot print_timing: id 1 | task 520 | prompt eval time = 88.94 ms / 4 tokens ( 22.23 ms per token, 44.97 tokens per second) 1.31.602.876 I slot print_timing: id 1 | task 520 | eval time = 8715.06 ms / 128 tokens ( 68.09 ms per token, 14.69 tokens per second) 1.31.602.877 I slot print_timing: id 1 | task 520 | total time = 8804.00 ms / 132 tokens 1.31.602.877 I slot print_timing: id 1 | task 520 | graphs reused = 1 1.31.602.880 I slot print_timing: id 1 | task 520 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.31.602.892 I statistics draft-mtp: #calls(b,g,a) = 115 511 7852, #gen drafts = 7866, #acc drafts = 5397, #gen tokens = 7866, #acc tokens = 5397, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.088, 878.526, 9.450 ms 1.31.602.921 I slot release: id 1 | task 520 | stop processing: n_tokens = 156, truncated = 0 1.31.608.247 I slot get_availabl: id 1 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.186 1.31.608.249 I srv get_availabl: updating prompt cache 1.31.608.291 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.31.608.295 I srv alloc: - prompt is already in the cache, skipping 1.31.608.296 I srv load: - looking for better prompt, base f_keep = 0.186, sim = 1.000 1.31.608.299 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.31.608.300 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.31.608.301 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.31.608.301 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.31.608.302 I srv get_availabl: prompt cache update took 0.05 ms 1.31.608.373 I slot launch_slot_: id 1 | task 628 | processing task, is_child = 0 1.31.608.374 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 1.31.608.390 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.31.608.391 I srv alloc: - prompt is already in the cache, skipping 1.31.611.326 I slot operator(): id 1 | task 628 | Checking checkpoint with [24, 24] against 28... 1.31.611.485 I srv operator(): Chat format: peg-native 1.31.615.191 W slot operator(): id 1 | task 628 | restored context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_past = 25, size = 63.009 MiB) 1.31.615.231 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.31.699.411 I slot print_timing: id 14 | task 541 | prompt eval time = 93.92 ms / 4 tokens ( 23.48 ms per token, 42.59 tokens per second) 1.31.699.414 I slot print_timing: id 14 | task 541 | eval time = 7155.06 ms / 128 tokens ( 55.90 ms per token, 17.89 tokens per second) 1.31.699.414 I slot print_timing: id 14 | task 541 | total time = 7248.98 ms / 132 tokens 1.31.699.415 I slot print_timing: id 14 | task 541 | graphs reused = 1 1.31.699.418 I slot print_timing: id 14 | task 541 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.31.699.432 I statistics draft-mtp: #calls(b,g,a) = 116 512 7866, #gen drafts = 7879, #acc drafts = 5405, #gen tokens = 7879, #acc tokens = 5405, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.089, 881.422, 9.466 ms 1.31.699.461 I slot release: id 14 | task 541 | stop processing: n_tokens = 155, truncated = 0 1.31.704.689 I slot get_availabl: id 14 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.31.704.693 I srv get_availabl: updating prompt cache 1.31.704.733 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.31.704.735 I srv alloc: - prompt is already in the cache, skipping 1.31.704.735 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.31.704.738 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.31.704.739 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.31.704.740 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.31.704.740 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.31.704.740 I srv get_availabl: prompt cache update took 0.05 ms 1.31.704.805 I slot launch_slot_: id 14 | task 630 | processing task, is_child = 0 1.31.704.806 I slot process_sing: id 8 | task -1 | saving idle slot to prompt cache 1.31.704.822 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.31.704.823 I srv alloc: - prompt is already in the cache, skipping 1.31.707.477 I slot operator(): id 14 | task 630 | Checking checkpoint with [23, 23] against 27... 1.31.708.612 I srv operator(): Chat format: peg-native 1.31.711.406 W slot operator(): id 14 | task 630 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.31.711.439 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.31.801.912 I slot print_timing: id 15 | task 560 | prompt eval time = 214.42 ms / 25 tokens ( 8.58 ms per token, 116.59 tokens per second) 1.31.801.915 I slot print_timing: id 15 | task 560 | eval time = 5678.10 ms / 114 tokens ( 49.81 ms per token, 20.08 tokens per second) 1.31.801.916 I slot print_timing: id 15 | task 560 | total time = 5892.52 ms / 139 tokens 1.31.801.917 I slot print_timing: id 15 | task 560 | graphs reused = 1 1.31.801.919 I slot print_timing: id 15 | task 560 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.31.801.936 I statistics draft-mtp: #calls(b,g,a) = 117 513 7893, #gen drafts = 7893, #acc drafts = 5422, #gen tokens = 7893, #acc tokens = 5422, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.090, 884.038, 9.501 ms 1.31.801.964 I slot release: id 15 | task 560 | stop processing: n_tokens = 139, truncated = 0 1.31.801.980 I slot get_availabl: id 8 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.31.801.981 I srv get_availabl: updating prompt cache 1.31.802.021 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.31.802.022 I srv alloc: - prompt is already in the cache, skipping 1.31.802.023 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.31.802.026 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.31.802.026 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.31.802.026 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.31.802.027 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.31.802.027 I srv get_availabl: prompt cache update took 0.05 ms 1.31.802.096 I slot launch_slot_: id 8 | task 632 | processing task, is_child = 0 1.31.802.096 I slot process_sing: id 15 | task -1 | saving idle slot to prompt cache 1.31.802.113 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.31.802.113 I srv alloc: - prompt is already in the cache, skipping 1.31.804.426 I slot operator(): id 8 | task 632 | Checking checkpoint with [20, 20] against 24... 1.31.808.346 W slot operator(): id 8 | task 632 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.31.808.395 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.31.809.925 I srv operator(): Chat format: peg-native 1.31.899.469 I slot get_availabl: id 15 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 1.31.899.472 I srv get_availabl: updating prompt cache 1.31.899.514 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.31.899.515 I srv alloc: - prompt is already in the cache, skipping 1.31.899.516 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 1.31.899.519 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.31.899.520 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.31.899.520 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.31.899.520 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.31.899.521 I srv get_availabl: prompt cache update took 0.05 ms 1.31.899.584 I slot launch_slot_: id 15 | task 634 | processing task, is_child = 0 1.31.901.032 I slot operator(): id 15 | task 634 | Checking checkpoint with [20, 20] against 3... 1.31.901.034 W slot operator(): id 15 | task 634 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.31.901.036 W slot operator(): id 15 | task 634 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.31.903.192 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.32.025.785 I slot create_check: id 15 | task 634 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.32.693.505 I slot print_timing: id 11 | task 556 | n_decoded = 101, tg = 14.86 t/s, tg_3s = 14.86 t/s 1.32.693.950 I slot print_timing: id 12 | task 579 | n_decoded = 101, tg = 20.24 t/s, tg_3s = 20.24 t/s 1.32.882.580 I slot print_timing: id 6 | task 576 | n_decoded = 100, tg = 18.52 t/s, tg_3s = 18.52 t/s 1.33.165.578 I slot print_timing: id 4 | task 559 | prompt eval time = 107.39 ms / 4 tokens ( 26.85 ms per token, 37.25 tokens per second) 1.33.165.581 I slot print_timing: id 4 | task 559 | eval time = 7152.71 ms / 128 tokens ( 55.88 ms per token, 17.90 tokens per second) 1.33.165.581 I slot print_timing: id 4 | task 559 | total time = 7260.11 ms / 132 tokens 1.33.165.582 I slot print_timing: id 4 | task 559 | graphs reused = 1 1.33.165.584 I slot print_timing: id 4 | task 559 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.33.165.608 I statistics draft-mtp: #calls(b,g,a) = 119 527 8097, #gen drafts = 8111, #acc drafts = 5564, #gen tokens = 8111, #acc tokens = 5564, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.093, 906.595, 9.763 ms 1.33.165.641 I slot release: id 4 | task 559 | stop processing: n_tokens = 155, truncated = 0 1.33.165.855 I slot print_timing: id 13 | task 539 | prompt eval time = 88.86 ms / 4 tokens ( 22.22 ms per token, 45.01 tokens per second) 1.33.165.855 I slot print_timing: id 13 | task 539 | eval time = 8723.19 ms / 128 tokens ( 68.15 ms per token, 14.67 tokens per second) 1.33.165.856 I slot print_timing: id 13 | task 539 | total time = 8812.06 ms / 132 tokens 1.33.165.856 I slot print_timing: id 13 | task 539 | graphs reused = 1 1.33.165.857 I slot print_timing: id 13 | task 539 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.33.165.860 I statistics draft-mtp: #calls(b,g,a) = 119 527 8097, #gen drafts = 8111, #acc drafts = 5564, #gen tokens = 8111, #acc tokens = 5564, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.093, 906.595, 9.763 ms 1.33.165.879 I slot release: id 13 | task 539 | stop processing: n_tokens = 156, truncated = 0 1.33.175.763 I srv operator(): Chat format: peg-native 1.33.175.772 I srv operator(): Chat format: peg-native 1.33.256.770 I slot get_availabl: id 4 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.33.256.772 I srv get_availabl: updating prompt cache 1.33.256.814 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.33.256.815 I srv alloc: - prompt is already in the cache, skipping 1.33.256.816 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.33.256.820 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.33.256.821 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.33.256.823 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.33.256.823 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.33.256.824 I srv get_availabl: prompt cache update took 0.05 ms 1.33.256.894 I slot launch_slot_: id 4 | task 649 | processing task, is_child = 0 1.33.256.895 I slot process_sing: id 13 | task -1 | saving idle slot to prompt cache 1.33.256.911 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.33.256.913 I srv alloc: - prompt is already in the cache, skipping 1.33.256.915 I slot get_availabl: id 13 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 1.33.256.916 I srv get_availabl: updating prompt cache 1.33.256.929 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.33.256.930 I srv alloc: - prompt is already in the cache, skipping 1.33.256.930 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 1.33.256.931 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.33.256.931 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.33.256.932 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.33.256.932 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.33.256.932 I srv get_availabl: prompt cache update took 0.02 ms 1.33.256.951 I slot launch_slot_: id 13 | task 650 | processing task, is_child = 0 1.33.259.710 I slot operator(): id 4 | task 649 | Checking checkpoint with [23, 23] against 3... 1.33.259.711 W slot operator(): id 4 | task 649 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.33.259.714 W slot operator(): id 4 | task 649 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.33.261.502 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.33.261.529 I slot operator(): id 13 | task 650 | Checking checkpoint with [24, 24] against 3... 1.33.261.530 W slot operator(): id 13 | task 650 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.33.261.532 W slot operator(): id 13 | task 650 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.33.263.837 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.33.380.870 I slot print_timing: id 12 | task 579 | prompt eval time = 212.14 ms / 25 tokens ( 8.49 ms per token, 117.84 tokens per second) 1.33.380.873 I slot print_timing: id 12 | task 579 | eval time = 5677.91 ms / 114 tokens ( 49.81 ms per token, 20.08 tokens per second) 1.33.380.874 I slot print_timing: id 12 | task 579 | total time = 5890.06 ms / 139 tokens 1.33.380.875 I slot print_timing: id 12 | task 579 | graphs reused = 1 1.33.380.877 I slot print_timing: id 12 | task 579 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.33.380.889 I statistics draft-mtp: #calls(b,g,a) = 119 529 8137, #gen drafts = 8139, #acc drafts = 5589, #gen tokens = 8139, #acc tokens = 5589, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.093, 912.050, 9.810 ms 1.33.380.918 I slot release: id 12 | task 579 | stop processing: n_tokens = 139, truncated = 0 1.33.389.366 I srv operator(): Chat format: peg-native 1.33.399.369 I slot create_check: id 4 | task 649 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.33.413.413 I slot create_check: id 13 | task 650 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.33.505.807 I slot get_availabl: id 12 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 1.33.505.809 I srv get_availabl: updating prompt cache 1.33.505.850 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.33.505.851 I srv alloc: - prompt is already in the cache, skipping 1.33.505.852 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 1.33.505.856 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.33.505.857 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.33.505.858 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.33.505.858 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.33.505.858 I srv get_availabl: prompt cache update took 0.05 ms 1.33.505.926 I slot launch_slot_: id 12 | task 653 | processing task, is_child = 0 1.33.507.923 I slot operator(): id 12 | task 653 | Checking checkpoint with [20, 20] against 3... 1.33.507.924 W slot operator(): id 12 | task 653 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.33.507.927 W slot operator(): id 12 | task 653 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.33.509.796 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.33.631.973 I slot create_check: id 12 | task 653 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.34.008.074 I slot print_timing: id 0 | task 574 | n_decoded = 101, tg = 14.98 t/s, tg_3s = 14.98 t/s 1.34.392.175 I slot print_timing: id 9 | task 599 | n_decoded = 101, tg = 20.25 t/s, tg_3s = 20.25 t/s 1.34.488.047 I slot print_timing: id 10 | task 595 | n_decoded = 100, tg = 18.54 t/s, tg_3s = 18.54 t/s 1.34.673.622 I slot print_timing: id 6 | task 576 | prompt eval time = 209.81 ms / 28 tokens ( 7.49 ms per token, 133.45 tokens per second) 1.34.673.625 I slot print_timing: id 6 | task 576 | eval time = 7190.34 ms / 128 tokens ( 56.17 ms per token, 17.80 tokens per second) 1.34.673.625 I slot print_timing: id 6 | task 576 | total time = 7400.15 ms / 156 tokens 1.34.673.626 I slot print_timing: id 6 | task 576 | graphs reused = 1 1.34.673.629 I slot print_timing: id 6 | task 576 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.34.673.645 I statistics draft-mtp: #calls(b,g,a) = 122 542 8326, #gen drafts = 8340, #acc drafts = 5715, #gen tokens = 8340, #acc tokens = 5715, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.095, 934.531, 10.038 ms 1.34.673.673 I slot release: id 6 | task 576 | stop processing: n_tokens = 155, truncated = 0 1.34.673.889 I slot print_timing: id 11 | task 556 | prompt eval time = 204.59 ms / 29 tokens ( 7.05 ms per token, 141.75 tokens per second) 1.34.673.889 I slot print_timing: id 11 | task 556 | eval time = 8775.77 ms / 128 tokens ( 68.56 ms per token, 14.59 tokens per second) 1.34.673.890 I slot print_timing: id 11 | task 556 | total time = 8980.36 ms / 157 tokens 1.34.673.890 I slot print_timing: id 11 | task 556 | graphs reused = 1 1.34.673.891 I slot print_timing: id 11 | task 556 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.34.673.893 I statistics draft-mtp: #calls(b,g,a) = 122 542 8326, #gen drafts = 8340, #acc drafts = 5715, #gen tokens = 8340, #acc tokens = 5715, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.095, 934.531, 10.038 ms 1.34.673.915 I slot release: id 11 | task 556 | stop processing: n_tokens = 156, truncated = 0 1.34.683.300 I srv operator(): Chat format: peg-native 1.34.683.748 I srv operator(): Chat format: peg-native 1.34.760.580 I slot print_timing: id 3 | task 584 | n_decoded = 101, tg = 15.19 t/s, tg_3s = 15.19 t/s 1.34.764.665 I slot get_availabl: id 6 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.181 1.34.764.667 I srv get_availabl: updating prompt cache 1.34.764.709 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.34.764.711 I srv alloc: - prompt is already in the cache, skipping 1.34.764.712 I srv load: - looking for better prompt, base f_keep = 0.181, sim = 1.000 1.34.764.715 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.34.764.715 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.34.764.716 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.34.764.717 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.34.764.717 I srv get_availabl: prompt cache update took 0.05 ms 1.34.764.785 I slot launch_slot_: id 6 | task 667 | processing task, is_child = 0 1.34.764.786 I slot process_sing: id 11 | task -1 | saving idle slot to prompt cache 1.34.764.802 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.34.764.803 I srv alloc: - prompt is already in the cache, skipping 1.34.764.804 I slot get_availabl: id 11 | task -1 | selected slot by LCP similarity, sim_best = 0.120 (> 0.100 thold), f_keep = 0.019 1.34.764.805 I srv get_availabl: updating prompt cache 1.34.764.818 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.34.764.818 I srv alloc: - prompt is already in the cache, skipping 1.34.764.819 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.120 1.34.764.819 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.34.764.820 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.34.764.820 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.34.764.820 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.34.764.820 I srv get_availabl: prompt cache update took 0.01 ms 1.34.764.837 I slot launch_slot_: id 11 | task 668 | processing task, is_child = 0 1.34.767.580 I slot operator(): id 6 | task 667 | Checking checkpoint with [23, 23] against 27... 1.34.771.463 W slot operator(): id 6 | task 667 | restored context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_past = 24, size = 63.001 MiB) 1.34.771.489 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.34.771.515 I slot operator(): id 11 | task 668 | Checking checkpoint with [24, 24] against 3... 1.34.771.516 W slot operator(): id 11 | task 668 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.34.771.517 W slot operator(): id 11 | task 668 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.34.773.172 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.34.896.405 I slot create_check: id 11 | task 668 | created context checkpoint 1 of 32 (pos_min = 20, pos_max = 20, n_tokens = 21, size = 62.978 MiB) 1.35.086.690 I slot print_timing: id 9 | task 599 | prompt eval time = 213.92 ms / 25 tokens ( 8.56 ms per token, 116.87 tokens per second) 1.35.086.693 I slot print_timing: id 9 | task 599 | eval time = 5681.26 ms / 114 tokens ( 49.84 ms per token, 20.07 tokens per second) 1.35.086.694 I slot print_timing: id 9 | task 599 | total time = 5895.18 ms / 139 tokens 1.35.086.695 I slot print_timing: id 9 | task 599 | graphs reused = 1 1.35.086.697 I slot print_timing: id 9 | task 599 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.35.086.711 I statistics draft-mtp: #calls(b,g,a) = 124 546 8393, #gen drafts = 8399, #acc drafts = 5764, #gen tokens = 8399, #acc tokens = 5764, #mean acc len = 1.69, #acc rate/pos = (0.687), dur(b,g,a) = 0.098, 943.500, 10.117 ms 1.35.086.739 I slot release: id 9 | task 599 | stop processing: n_tokens = 139, truncated = 0 1.35.095.134 I srv operator(): Chat format: peg-native 1.35.179.974 I slot get_availabl: id 9 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.022 1.35.179.976 I srv get_availabl: updating prompt cache 1.35.180.017 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.35.180.018 I srv alloc: - prompt is already in the cache, skipping 1.35.180.019 I srv load: - looking for better prompt, base f_keep = 0.022, sim = 0.103 1.35.180.022 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.35.180.023 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.35.180.023 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.35.180.024 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.35.180.025 I srv get_availabl: prompt cache update took 0.05 ms 1.35.180.092 I slot launch_slot_: id 9 | task 673 | processing task, is_child = 0 1.35.182.460 I slot operator(): id 9 | task 673 | Checking checkpoint with [20, 20] against 3... 1.35.182.461 W slot operator(): id 9 | task 673 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.35.182.464 W slot operator(): id 9 | task 673 | erased invalidated context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_swa = 0, pos_next = 0, size = 62.978 MiB) 1.35.184.289 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.35.306.616 I slot create_check: id 9 | task 673 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.35.493.073 I slot print_timing: id 5 | task 614 | n_decoded = 101, tg = 20.07 t/s, tg_3s = 20.07 t/s 1.35.779.845 I slot print_timing: id 7 | task 612 | n_decoded = 100, tg = 18.44 t/s, tg_3s = 18.44 t/s 1.35.966.186 I slot print_timing: id 0 | task 574 | prompt eval time = 88.83 ms / 4 tokens ( 22.21 ms per token, 45.03 tokens per second) 1.35.966.189 I slot print_timing: id 0 | task 574 | eval time = 8700.82 ms / 128 tokens ( 67.98 ms per token, 14.71 tokens per second) 1.35.966.189 I slot print_timing: id 0 | task 574 | total time = 8789.65 ms / 132 tokens 1.35.966.190 I slot print_timing: id 0 | task 574 | graphs reused = 1 1.35.966.192 I slot print_timing: id 0 | task 574 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.35.966.208 I statistics draft-mtp: #calls(b,g,a) = 125 555 8524, #gen drafts = 8539, #acc drafts = 5848, #gen tokens = 8539, #acc tokens = 5848, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.098, 958.808, 10.267 ms 1.35.966.237 I slot release: id 0 | task 574 | stop processing: n_tokens = 156, truncated = 0 1.35.975.184 I srv operator(): Chat format: peg-native 1.36.062.328 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, sim_best = 0.107 (> 0.100 thold), f_keep = 0.019 1.36.062.331 I srv get_availabl: updating prompt cache 1.36.062.373 W srv prompt_save: - saving prompt with length 156, total state size = 67.085 MiB (draft: 1.222 MiB) 1.36.062.375 I srv alloc: - prompt is already in the cache, skipping 1.36.062.376 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.107 1.36.062.379 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.36.062.380 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.36.062.380 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.36.062.381 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.36.062.381 I srv get_availabl: prompt cache update took 0.05 ms 1.36.062.450 I slot launch_slot_: id 0 | task 683 | processing task, is_child = 0 1.36.063.903 I slot operator(): id 0 | task 683 | Checking checkpoint with [24, 24] against 3... 1.36.063.904 W slot operator(): id 0 | task 683 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.36.063.906 W slot operator(): id 0 | task 683 | erased invalidated context checkpoint (pos_min = 24, pos_max = 24, n_tokens = 25, n_swa = 0, pos_next = 0, size = 63.009 MiB) 1.36.065.889 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.36.167.448 I slot print_timing: id 5 | task 614 | prompt eval time = 95.26 ms / 4 tokens ( 23.82 ms per token, 41.99 tokens per second) 1.36.167.451 I slot print_timing: id 5 | task 614 | eval time = 5707.83 ms / 114 tokens ( 50.07 ms per token, 19.97 tokens per second) 1.36.167.451 I slot print_timing: id 5 | task 614 | total time = 5803.09 ms / 118 tokens 1.36.167.452 I slot print_timing: id 5 | task 614 | graphs reused = 1 1.36.167.454 I slot print_timing: id 5 | task 614 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.36.167.469 I statistics draft-mtp: #calls(b,g,a) = 125 557 8559, #gen drafts = 8569, #acc drafts = 5874, #gen tokens = 8569, #acc tokens = 5874, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.098, 961.633, 10.310 ms 1.36.167.497 I slot release: id 5 | task 614 | stop processing: n_tokens = 139, truncated = 0 1.36.176.524 I srv operator(): Chat format: peg-native 1.36.187.971 I slot create_check: id 0 | task 683 | created context checkpoint 1 of 32 (pos_min = 23, pos_max = 23, n_tokens = 24, size = 63.001 MiB) 1.36.272.644 I slot print_timing: id 10 | task 595 | prompt eval time = 212.92 ms / 28 tokens ( 7.60 ms per token, 131.51 tokens per second) 1.36.272.647 I slot print_timing: id 10 | task 595 | eval time = 7179.73 ms / 128 tokens ( 56.09 ms per token, 17.83 tokens per second) 1.36.272.647 I slot print_timing: id 10 | task 595 | total time = 7392.65 ms / 156 tokens 1.36.272.648 I slot print_timing: id 10 | task 595 | graphs reused = 1 1.36.272.651 I slot print_timing: id 10 | task 595 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.36.272.664 I statistics draft-mtp: #calls(b,g,a) = 126 558 8569, #gen drafts = 8582, #acc drafts = 5882, #gen tokens = 8582, #acc tokens = 5882, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.099, 964.277, 10.322 ms 1.36.272.691 I slot release: id 10 | task 595 | stop processing: n_tokens = 155, truncated = 0 1.36.277.608 I slot get_availabl: id 5 | task -1 | selected slot by LCP similarity, sim_best = 1.000 (> 0.100 thold), f_keep = 0.180 1.36.277.612 I srv get_availabl: updating prompt cache 1.36.277.652 W srv prompt_save: - saving prompt with length 139, total state size = 66.620 MiB (draft: 1.089 MiB) 1.36.277.654 I srv alloc: - prompt is already in the cache, skipping 1.36.277.654 I srv load: - looking for better prompt, base f_keep = 0.180, sim = 1.000 1.36.277.658 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.36.277.659 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.36.277.659 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.36.277.659 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.36.277.660 I srv get_availabl: prompt cache update took 0.05 ms 1.36.277.723 I slot launch_slot_: id 5 | task 686 | processing task, is_child = 0 1.36.277.724 I slot process_sing: id 10 | task -1 | saving idle slot to prompt cache 1.36.277.739 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.36.277.740 I srv alloc: - prompt is already in the cache, skipping 1.36.280.416 I slot operator(): id 5 | task 686 | Checking checkpoint with [20, 20] against 24... 1.36.281.639 I srv operator(): Chat format: peg-native 1.36.284.310 W slot operator(): id 5 | task 686 | restored context checkpoint (pos_min = 20, pos_max = 20, n_tokens = 21, n_past = 21, size = 62.978 MiB) 1.36.284.347 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.36.375.004 I slot get_availabl: id 10 | task -1 | selected slot by LCP similarity, sim_best = 0.103 (> 0.100 thold), f_keep = 0.019 1.36.375.006 I srv get_availabl: updating prompt cache 1.36.375.047 W srv prompt_save: - saving prompt with length 155, total state size = 67.058 MiB (draft: 1.214 MiB) 1.36.375.049 I srv alloc: - prompt is already in the cache, skipping 1.36.375.050 I srv load: - looking for better prompt, base f_keep = 0.019, sim = 0.103 1.36.375.053 I srv update: - cache state: 3 prompts, 389.751 MiB (limits: 8192.000 MiB, 8192 tokens, 9458 est) 1.36.375.054 I srv update: - prompt 0x6328042bc260: 139 tokens, checkpoints: 1, 129.598 MiB 1.36.375.054 I srv update: - prompt 0x6327fbdda650: 155 tokens, checkpoints: 1, 130.059 MiB 1.36.375.055 I srv update: - prompt 0x6328042bbe70: 156 tokens, checkpoints: 1, 130.094 MiB 1.36.375.055 I srv get_availabl: prompt cache update took 0.05 ms 1.36.375.124 I slot launch_slot_: id 10 | task 688 | processing task, is_child = 0 1.36.377.180 I slot operator(): id 10 | task 688 | Checking checkpoint with [23, 23] against 3... 1.36.377.182 W slot operator(): id 10 | task 688 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 1.36.377.184 W slot operator(): id 10 | task 688 | erased invalidated context checkpoint (pos_min = 23, pos_max = 23, n_tokens = 24, n_swa = 0, pos_next = 0, size = 63.001 MiB) 1.36.379.016 I srv stream_sessi: stream_session_attach_pipe: conv_id= (empty=1) 1.36.502.099 I slot create_check: id 10 | task 688 | created context checkpoint 1 of 32 (pos_min = 24, pos_max = 24, n_tokens = 25, size = 63.009 MiB) 1.36.781.918 I slot print_timing: id 3 | task 584 | prompt eval time = 214.35 ms / 29 tokens ( 7.39 ms per token, 135.29 tokens per second) 1.36.781.921 I slot print_timing: id 3 | task 584 | eval time = 8669.71 ms / 128 tokens ( 67.73 ms per token, 14.76 tokens per second) 1.36.781.922 I slot print_timing: id 3 | task 584 | total time = 8884.06 ms / 157 tokens 1.36.781.923 I slot print_timing: id 3 | task 584 | graphs reused = 1 1.36.781.925 I slot print_timing: id 3 | task 584 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.36.781.940 I statistics draft-mtp: #calls(b,g,a) = 128 563 8642, #gen drafts = 8657, #acc drafts = 5931, #gen tokens = 8657, #acc tokens = 5931, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.101, 974.812, 10.418 ms 1.36.781.971 I slot release: id 3 | task 584 | stop processing: n_tokens = 156, truncated = 0 1.36.963.814 I slot print_timing: id 2 | task 610 | n_decoded = 101, tg = 15.06 t/s, tg_3s = 15.06 t/s 1.36.965.987 I slot print_timing: id 8 | task 632 | n_decoded = 101, tg = 19.91 t/s, tg_3s = 19.91 t/s 1.37.238.092 I slot print_timing: id 14 | task 630 | n_decoded = 100, tg = 18.38 t/s, tg_3s = 18.38 t/s 1.37.504.335 I slot print_timing: id 7 | task 612 | prompt eval time = 89.01 ms / 4 tokens ( 22.25 ms per token, 44.94 tokens per second) 1.37.504.338 I slot print_timing: id 7 | task 612 | eval time = 7147.60 ms / 128 tokens ( 55.84 ms per token, 17.91 tokens per second) 1.37.504.339 I slot print_timing: id 7 | task 612 | total time = 7236.61 ms / 132 tokens 1.37.504.340 I slot print_timing: id 7 | task 612 | graphs reused = 1 1.37.504.343 I slot print_timing: id 7 | task 612 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.37.504.360 I statistics draft-mtp: #calls(b,g,a) = 128 571 8762, #gen drafts = 8776, #acc drafts = 6011, #gen tokens = 8776, #acc tokens = 6011, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.101, 991.344, 10.562 ms 1.37.504.389 I slot release: id 7 | task 612 | stop processing: n_tokens = 155, truncated = 0 1.37.593.122 I slot print_timing: id 8 | task 632 | prompt eval time = 89.09 ms / 4 tokens ( 22.27 ms per token, 44.90 tokens per second) 1.37.593.125 I slot print_timing: id 8 | task 632 | eval time = 5699.59 ms / 114 tokens ( 50.00 ms per token, 20.00 tokens per second) 1.37.593.125 I slot print_timing: id 8 | task 632 | total time = 5788.68 ms / 118 tokens 1.37.593.126 I slot print_timing: id 8 | task 632 | graphs reused = 1 1.37.593.129 I slot print_timing: id 8 | task 632 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.37.593.146 I statistics draft-mtp: #calls(b,g,a) = 128 572 8783, #gen drafts = 8790, #acc drafts = 6025, #gen tokens = 8790, #acc tokens = 6025, #mean acc len = 1.69, #acc rate/pos = (0.686), dur(b,g,a) = 0.101, 993.988, 10.589 ms 1.37.593.171 I slot release: id 8 | task 632 | stop processing: n_tokens = 139, truncated = 0 1.38.306.390 I slot print_timing: id 1 | task 628 | n_decoded = 101, tg = 15.29 t/s, tg_3s = 15.29 t/s 1.38.307.264 I slot print_timing: id 4 | task 649 | n_decoded = 101, tg = 21.01 t/s, tg_3s = 21.01 t/s 1.38.548.415 I slot print_timing: id 15 | task 634 | n_decoded = 101, tg = 15.70 t/s, tg_3s = 15.70 t/s 1.38.621.687 I slot print_timing: id 2 | task 610 | prompt eval time = 88.15 ms / 4 tokens ( 22.04 ms per token, 45.38 tokens per second) 1.38.621.689 I slot print_timing: id 2 | task 610 | eval time = 8362.38 ms / 128 tokens ( 65.33 ms per token, 15.31 tokens per second) 1.38.621.689 I slot print_timing: id 2 | task 610 | total time = 8450.53 ms / 132 tokens 1.38.621.690 I slot print_timing: id 2 | task 610 | graphs reused = 1 1.38.621.692 I slot print_timing: id 2 | task 610 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.38.621.706 I statistics draft-mtp: #calls(b,g,a) = 128 585 8946, #gen drafts = 8958, #acc drafts = 6132, #gen tokens = 8958, #acc tokens = 6132, #mean acc len = 1.69, #acc rate/pos = (0.685), dur(b,g,a) = 0.101, 1026.146, 10.785 ms 1.38.621.734 I slot release: id 2 | task 610 | stop processing: n_tokens = 156, truncated = 0 1.38.625.565 I slot print_timing: id 13 | task 650 | n_decoded = 100, tg = 19.51 t/s, tg_3s = 19.51 t/s 1.38.694.381 I slot print_timing: id 14 | task 630 | prompt eval time = 89.06 ms / 4 tokens ( 22.27 ms per token, 44.91 tokens per second) 1.38.694.383 I slot print_timing: id 14 | task 630 | eval time = 6897.82 ms / 128 tokens ( 53.89 ms per token, 18.56 tokens per second) 1.38.694.384 I slot print_timing: id 14 | task 630 | total time = 6986.89 ms / 132 tokens 1.38.694.384 I slot print_timing: id 14 | task 630 | graphs reused = 1 1.38.694.387 I slot print_timing: id 14 | task 630 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.38.694.404 I statistics draft-mtp: #calls(b,g,a) = 128 586 8958, #gen drafts = 8969, #acc drafts = 6138, #gen tokens = 8969, #acc tokens = 6138, #mean acc len = 1.69, #acc rate/pos = (0.685), dur(b,g,a) = 0.101, 1028.807, 10.798 ms 1.38.694.425 I slot release: id 14 | task 630 | stop processing: n_tokens = 155, truncated = 0 1.38.832.815 I slot print_timing: id 4 | task 649 | prompt eval time = 240.71 ms / 25 tokens ( 9.63 ms per token, 103.86 tokens per second) 1.38.832.818 I slot print_timing: id 4 | task 649 | eval time = 5332.38 ms / 114 tokens ( 46.78 ms per token, 21.38 tokens per second) 1.38.832.818 I slot print_timing: id 4 | task 649 | total time = 5573.09 ms / 139 tokens 1.38.832.819 I slot print_timing: id 4 | task 649 | graphs reused = 1 1.38.832.821 I slot print_timing: id 4 | task 649 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.38.832.834 I statistics draft-mtp: #calls(b,g,a) = 128 588 8983, #gen drafts = 8991, #acc drafts = 6153, #gen tokens = 8991, #acc tokens = 6153, #mean acc len = 1.68, #acc rate/pos = (0.685), dur(b,g,a) = 0.101, 1034.082, 10.830 ms 1.38.832.855 I slot release: id 4 | task 649 | stop processing: n_tokens = 139, truncated = 0 1.39.333.188 I slot print_timing: id 11 | task 668 | n_decoded = 101, tg = 23.24 t/s, tg_3s = 23.24 t/s 1.39.518.760 I slot print_timing: id 6 | task 667 | n_decoded = 100, tg = 21.53 t/s, tg_3s = 21.53 t/s 1.39.582.157 I slot print_timing: id 12 | task 653 | n_decoded = 101, tg = 17.24 t/s, tg_3s = 17.24 t/s 1.39.640.628 I slot print_timing: id 1 | task 628 | prompt eval time = 87.86 ms / 4 tokens ( 21.96 ms per token, 45.53 tokens per second) 1.39.640.631 I slot print_timing: id 1 | task 628 | eval time = 7941.42 ms / 128 tokens ( 62.04 ms per token, 16.12 tokens per second) 1.39.640.631 I slot print_timing: id 1 | task 628 | total time = 8029.28 ms / 132 tokens 1.39.640.632 I slot print_timing: id 1 | task 628 | graphs reused = 1 1.39.640.635 I slot print_timing: id 1 | task 628 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.39.640.650 I statistics draft-mtp: #calls(b,g,a) = 128 601 9111, #gen drafts = 9120, #acc drafts = 6233, #gen tokens = 9120, #acc tokens = 6233, #mean acc len = 1.68, #acc rate/pos = (0.684), dur(b,g,a) = 0.101, 1066.823, 10.980 ms 1.39.640.677 I slot release: id 1 | task 628 | stop processing: n_tokens = 156, truncated = 0 1.39.752.378 I slot print_timing: id 13 | task 650 | prompt eval time = 239.16 ms / 28 tokens ( 8.54 ms per token, 117.08 tokens per second) 1.39.752.380 I slot print_timing: id 13 | task 650 | eval time = 6251.67 ms / 128 tokens ( 48.84 ms per token, 20.47 tokens per second) 1.39.752.380 I slot print_timing: id 13 | task 650 | total time = 6490.83 ms / 156 tokens 1.39.752.381 I slot print_timing: id 13 | task 650 | graphs reused = 1 1.39.752.384 I slot print_timing: id 13 | task 650 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.39.752.397 I statistics draft-mtp: #calls(b,g,a) = 128 603 9129, #gen drafts = 9137, #acc drafts = 6243, #gen tokens = 9137, #acc tokens = 6243, #mean acc len = 1.68, #acc rate/pos = (0.684), dur(b,g,a) = 0.101, 1071.834, 11.005 ms 1.39.752.419 I slot release: id 13 | task 650 | stop processing: n_tokens = 155, truncated = 0 1.39.754.822 I slot print_timing: id 11 | task 668 | prompt eval time = 215.31 ms / 25 tokens ( 8.61 ms per token, 116.11 tokens per second) 1.39.754.823 I slot print_timing: id 11 | task 668 | eval time = 4767.98 ms / 114 tokens ( 41.82 ms per token, 23.91 tokens per second) 1.39.754.824 I slot print_timing: id 11 | task 668 | total time = 4983.30 ms / 139 tokens 1.39.754.824 I slot print_timing: id 11 | task 668 | graphs reused = 1 1.39.754.826 I slot print_timing: id 11 | task 668 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.39.754.833 I statistics draft-mtp: #calls(b,g,a) = 128 603 9135, #gen drafts = 9137, #acc drafts = 6247, #gen tokens = 9137, #acc tokens = 6247, #mean acc len = 1.68, #acc rate/pos = (0.684), dur(b,g,a) = 0.101, 1071.834, 11.014 ms 1.39.754.850 I slot release: id 11 | task 668 | stop processing: n_tokens = 139, truncated = 0 1.39.796.386 I slot print_timing: id 15 | task 634 | prompt eval time = 214.44 ms / 29 tokens ( 7.39 ms per token, 135.24 tokens per second) 1.39.796.389 I slot print_timing: id 15 | task 634 | eval time = 7680.90 ms / 128 tokens ( 60.01 ms per token, 16.66 tokens per second) 1.39.796.389 I slot print_timing: id 15 | task 634 | total time = 7895.34 ms / 157 tokens 1.39.796.390 I slot print_timing: id 15 | task 634 | graphs reused = 1 1.39.796.392 I slot print_timing: id 15 | task 634 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.39.796.407 I statistics draft-mtp: #calls(b,g,a) = 128 604 9137, #gen drafts = 9143, #acc drafts = 6248, #gen tokens = 9143, #acc tokens = 6248, #mean acc len = 1.68, #acc rate/pos = (0.684), dur(b,g,a) = 0.101, 1074.005, 11.016 ms 1.39.796.428 I slot release: id 15 | task 634 | stop processing: n_tokens = 156, truncated = 0 1.40.029.486 I slot print_timing: id 5 | task 686 | n_decoded = 101, tg = 27.59 t/s, tg_3s = 27.59 t/s 1.40.145.341 I slot print_timing: id 0 | task 683 | n_decoded = 100, tg = 25.82 t/s, tg_3s = 25.82 t/s 1.40.297.426 I slot print_timing: id 6 | task 667 | prompt eval time = 107.41 ms / 4 tokens ( 26.85 ms per token, 37.24 tokens per second) 1.40.297.429 I slot print_timing: id 6 | task 667 | eval time = 5422.42 ms / 128 tokens ( 42.36 ms per token, 23.61 tokens per second) 1.40.297.429 I slot print_timing: id 6 | task 667 | total time = 5529.83 ms / 132 tokens 1.40.297.430 I slot print_timing: id 6 | task 667 | graphs reused = 1 1.40.297.432 I slot print_timing: id 6 | task 667 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.40.297.447 I statistics draft-mtp: #calls(b,g,a) = 128 617 9215, #gen drafts = 9220, #acc drafts = 6294, #gen tokens = 9220, #acc tokens = 6294, #mean acc len = 1.68, #acc rate/pos = (0.683), dur(b,g,a) = 0.101, 1101.543, 11.112 ms 1.40.297.468 I slot release: id 6 | task 667 | stop processing: n_tokens = 155, truncated = 0 1.40.298.414 I slot print_timing: id 5 | task 686 | prompt eval time = 88.94 ms / 4 tokens ( 22.23 ms per token, 44.98 tokens per second) 1.40.298.415 I slot print_timing: id 5 | task 686 | eval time = 3929.05 ms / 114 tokens ( 34.47 ms per token, 29.01 tokens per second) 1.40.298.415 I slot print_timing: id 5 | task 686 | total time = 4017.99 ms / 118 tokens 1.40.298.416 I slot print_timing: id 5 | task 686 | graphs reused = 1 1.40.298.418 I slot print_timing: id 5 | task 686 | draft acceptance = 0.96552 ( 56 accepted / 58 generated), mean acceptance length = 1.97, acceptance rate per position = (0.966) 1.40.298.424 I statistics draft-mtp: #calls(b,g,a) = 128 617 9217, #gen drafts = 9220, #acc drafts = 6296, #gen tokens = 9220, #acc tokens = 6296, #mean acc len = 1.68, #acc rate/pos = (0.683), dur(b,g,a) = 0.101, 1101.543, 11.116 ms 1.40.298.438 I slot release: id 5 | task 686 | stop processing: n_tokens = 139, truncated = 0 1.40.298.904 I slot print_timing: id 9 | task 673 | n_decoded = 101, tg = 20.60 t/s, tg_3s = 20.60 t/s 1.40.374.991 I slot print_timing: id 12 | task 653 | prompt eval time = 214.07 ms / 29 tokens ( 7.38 ms per token, 135.47 tokens per second) 1.40.374.994 I slot print_timing: id 12 | task 653 | eval time = 6652.99 ms / 128 tokens ( 51.98 ms per token, 19.24 tokens per second) 1.40.374.994 I slot print_timing: id 12 | task 653 | total time = 6867.05 ms / 157 tokens 1.40.374.995 I slot print_timing: id 12 | task 653 | graphs reused = 1 1.40.374.997 I slot print_timing: id 12 | task 653 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.40.375.008 I statistics draft-mtp: #calls(b,g,a) = 128 620 9228, #gen drafts = 9231, #acc drafts = 6301, #gen tokens = 9231, #acc tokens = 6301, #mean acc len = 1.68, #acc rate/pos = (0.683), dur(b,g,a) = 0.101, 1105.765, 11.128 ms 1.40.375.027 I slot release: id 12 | task 653 | stop processing: n_tokens = 156, truncated = 0 1.40.572.139 I slot print_timing: id 10 | task 688 | n_decoded = 101, tg = 25.38 t/s, tg_3s = 25.38 t/s 1.40.589.098 I slot print_timing: id 0 | task 683 | prompt eval time = 208.51 ms / 28 tokens ( 7.45 ms per token, 134.29 tokens per second) 1.40.589.100 I slot print_timing: id 0 | task 683 | eval time = 4316.68 ms / 128 tokens ( 33.72 ms per token, 29.65 tokens per second) 1.40.589.101 I slot print_timing: id 0 | task 683 | total time = 4525.19 ms / 156 tokens 1.40.589.101 I slot print_timing: id 0 | task 683 | graphs reused = 1 1.40.589.104 I slot print_timing: id 0 | task 683 | draft acceptance = 0.75000 ( 54 accepted / 72 generated), mean acceptance length = 1.75, acceptance rate per position = (0.750) 1.40.589.115 I statistics draft-mtp: #calls(b,g,a) = 128 631 9261, #gen drafts = 9263, #acc drafts = 6317, #gen tokens = 9263, #acc tokens = 6317, #mean acc len = 1.68, #acc rate/pos = (0.682), dur(b,g,a) = 0.101, 1116.769, 11.160 ms 1.40.589.133 I slot release: id 0 | task 683 | stop processing: n_tokens = 155, truncated = 0 1.40.665.736 I slot print_timing: id 9 | task 673 | prompt eval time = 214.14 ms / 29 tokens ( 7.38 ms per token, 135.42 tokens per second) 1.40.665.738 I slot print_timing: id 9 | task 673 | eval time = 5269.13 ms / 128 tokens ( 41.17 ms per token, 24.29 tokens per second) 1.40.665.739 I slot print_timing: id 9 | task 673 | total time = 5483.27 ms / 157 tokens 1.40.665.739 I slot print_timing: id 9 | task 673 | graphs reused = 1 1.40.665.741 I slot print_timing: id 9 | task 673 | draft acceptance = 0.44828 ( 39 accepted / 87 generated), mean acceptance length = 1.45, acceptance rate per position = (0.448) 1.40.665.752 I statistics draft-mtp: #calls(b,g,a) = 128 637 9273, #gen drafts = 9274, #acc drafts = 6320, #gen tokens = 9274, #acc tokens = 6320, #mean acc len = 1.68, #acc rate/pos = (0.682), dur(b,g,a) = 0.101, 1120.083, 11.172 ms 1.40.665.767 I slot release: id 9 | task 673 | stop processing: n_tokens = 156, truncated = 0 1.40.746.939 I slot print_timing: id 10 | task 688 | prompt eval time = 214.94 ms / 29 tokens ( 7.41 ms per token, 134.92 tokens per second) 1.40.746.942 I slot print_timing: id 10 | task 688 | eval time = 4154.81 ms / 128 tokens ( 32.46 ms per token, 30.81 tokens per second) 1.40.746.942 I slot print_timing: id 10 | task 688 | total time = 4369.75 ms / 157 tokens 1.40.746.943 I slot print_timing: id 10 | task 688 | graphs reused = 13 1.40.746.944 I slot print_timing: id 10 | task 688 | draft acceptance = 0.43182 ( 38 accepted / 88 generated), mean acceptance length = 1.43, acceptance rate per position = (0.432) 1.40.746.955 I statistics draft-mtp: #calls(b,g,a) = 128 649 9286, #gen drafts = 9286, #acc drafts = 6326, #gen tokens = 9286, #acc tokens = 6326, #mean acc len = 1.68, #acc rate/pos = (0.681), dur(b,g,a) = 0.101, 1125.711, 11.183 ms 1.40.746.968 I slot release: id 10 | task 688 | stop processing: n_tokens = 156, truncated = 0 1.40.746.977 I srv update_slots: all slots are idle