tacos4me commited on
Commit
6bd1d23
·
verified ·
1 Parent(s): 30ff923

evidence: TF@8666 dump, 8666@512k confirmation log, convert_fp8visual.py

Browse files
scripts/confirm-8666-512k.log ADDED
@@ -0,0 +1,485 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0
 
1
 
 
 
 
 
2
 
3
 
4
 
5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [budget] waiting for GPU lock...
2
+ [budget] holding GPU lock
3
+ (APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345]
4
+ (APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345] █ █ █▄ ▄█
5
+ (APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.1.dev20051+g487ecf187
6
+ (APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345] █▄█▀ █ █ █ █ model /model
7
+ (APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
8
+ (APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345]
9
+ (APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:273] non-default args: {'model_tag': '/model', 'enable_auto_tool_choice': True, 'tool_call_parser': 'glm47', 'host': '0.0.0.0', 'port': 18998, 'model': '/model', 'trust_remote_code': True, 'max_model_len': 524288, 'served_model_name': ['glm-max'], 'reasoning_parser': 'glm45', 'tensor_parallel_size': 2, 'gpu_memory_utilization': 0.95, 'kv_cache_memory_bytes': 3484000000, 'kv_cache_dtype': 'fp8', 'enable_prefix_caching': True, 'limit_mm_per_prompt': {'image': 1, 'video': 0}, 'max_num_batched_tokens': 1024, 'max_num_seqs': 1, 'enable_flashinfer_autotune': False, 'speculative_config': {'method': 'mtp', 'num_speculative_tokens': 1}, 'compilation_config': {'mode': None, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': [], 'ir_enable_torch_wrap': None, 'splitting_ops': None, 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': None, 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': None, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [1, 2], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': None, 'pass_config': {}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': None, 'static_all_moe_layers': []}, 'kernel_config': KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=[], fused_add_rms_norm=[]), enable_flashinfer_autotune=None, enable_cutedsl_warmup=False, enable_jit_warmup=False, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')}
10
+ (APIServer pid=59) WARNING 08-30 23:38:57 [envs.py:2187] Unknown vLLM environment variable detected: VLLM_KVQ_TILES
11
+ (APIServer pid=59) INFO 08-30 23:38:57 [model.py:679] Resolved architecture: Glm5NextForConditionalGeneration
12
+ (APIServer pid=59) INFO 08-30 23:38:57 [model.py:1972] Using max model len 524288
13
+ (APIServer pid=59) INFO 08-30 23:39:04 [cache.py:282] Using fp8 data type to store kv cache. It reduces the GPU memory footprint and boosts the performance. Meanwhile, it may cause accuracy drop without a proper scaling factor
14
+ (APIServer pid=59) INFO 08-30 23:39:04 [model.py:679] Resolved architecture: Glm5NextMTPModel
15
+ (APIServer pid=59) INFO 08-30 23:39:04 [model.py:1972] Using max model len 1048576
16
+ (APIServer pid=59) INFO 08-30 23:39:04 [speculative.py:1256] Overriding draft model max model len from 1048576 to 524288
17
+ (APIServer pid=59) INFO 08-30 23:39:04 [scheduler.py:242] Chunked prefill is enabled with max_num_batched_tokens=1024.
18
+ (APIServer pid=59) INFO 08-30 23:39:04 [config.py:605] Mamba cache mode is set to 'align' for Glm5NextForConditionalGeneration by default when prefix caching is enabled
19
+ (APIServer pid=59) WARNING 08-30 23:39:04 [modelopt.py:379] Detected ModelOpt fp8 checkpoint (quant_algo=FP8). Please note that the format is experimental and could change.
20
+ (APIServer pid=59) WARNING 08-30 23:39:04 [modelopt.py:1012] Detected ModelOpt NVFP4 checkpoint (quant_algo=NVFP4). Please note that the format is experimental and could change in future.
21
+ (APIServer pid=59) WARNING 08-30 23:39:04 [modelopt.py:1012] Detected ModelOpt NVFP4 checkpoint (quant_algo=W4A16_NVFP4). Please note that the format is experimental and could change in future.
22
+ (APIServer pid=59) WARNING 08-30 23:39:04 [modelopt.py:1676] Detected ModelOpt MXFP8 checkpoint. Please note that the format is experimental and could change in future.
23
+ (APIServer pid=59) INFO 08-30 23:39:04 [vllm.py:1319] Auto-enabling VLLM_USE_BREAKABLE_CUDAGRAPH=1. Set VLLM_USE_BREAKABLE_CUDAGRAPH=0 to opt out.
24
+ (APIServer pid=59) INFO 08-30 23:39:04 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
25
+ (APIServer pid=59) WARNING 08-30 23:39:04 [vllm.py:1849] max_num_scheduled_tokens is set to 1024 based on the speculative decoding settings. This may lead to suboptimal performance. Consider increasing max_num_batched_tokens to accommodate the additional draft token slots, or decrease num_speculative_tokens.
26
+ (APIServer pid=59) INFO 08-30 23:39:04 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
27
+ (APIServer pid=59) [transformers] The following generation flags are not valid and may be ignored: ['top_p']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
28
+ (EngineCore pid=327) INFO 08-30 23:39:12 [core.py:122] Initializing a V1 LLM engine (v0.1.dev20051+g487ecf187) with config: model='/model', speculative_config=SpeculativeConfig(method='mtp', model='/model', num_spec_tokens=1), tokenizer='/model', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=524288, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=modelopt_mixed, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='glm45', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=glm-max, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [1024], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.FULL_AND_PIECEWISE: (2, 1)>, 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 2, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=False, enable_cutedsl_warmup=False, enable_jit_warmup=False, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')
29
+ (EngineCore pid=327) INFO 08-30 23:39:12 [multiproc_executor.py:149] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=192.168.1.80 (local), world_size=2, local_world_size=2
30
+ (Worker pid=461) INFO 08-30 23:39:20 [parallel_state.py:1638] world_size=2 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_ff8e5afc5b1143acad305c975dd061fc backend=nccl
31
+ (Worker pid=612) INFO 08-30 23:39:24 [parallel_state.py:1638] world_size=2 rank=1 local_rank=1 distributed_init_method=file:///tmp/vllm_dist_ff8e5afc5b1143acad305c975dd061fc backend=nccl
32
+ (Worker pid=461) INFO 08-30 23:39:24 [pynccl.py:113] vLLM is using nccl==2.30.7
33
+ (Worker pid=612) WARNING 08-30 23:39:25 [symm_mem.py:67] SymmMemCommunicator: Device capability 12.0 not supported, communicator is not available.
34
+ (Worker pid=461) WARNING 08-30 23:39:25 [symm_mem.py:67] SymmMemCommunicator: Device capability 12.0 not supported, communicator is not available.
35
+ (Worker pid=461) INFO 08-30 23:39:25 [cuda_communicator.py:266] Using ['CUSTOM', 'PYNCCL'] all-reduce backends (in dispatch order) for group 'tp:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL'].
36
+ (Worker pid=461) INFO 08-30 23:39:25 [cuda_communicator.py:266] Using ['PYNCCL'] all-reduce backends (in dispatch order) for group 'ep:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL'].
37
+ (Worker pid=461) INFO 08-30 23:39:25 [parallel_state.py:1982] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
38
+ (Worker pid=461) INFO 08-30 23:39:25 [gpu_worker.py:386] Using V2 Model Runner
39
+ (Worker_TP0 pid=461) INFO 08-30 23:39:25 [model_runner.py:353] Loading model from scratch...
40
+ (Worker_TP0 pid=461) INFO 08-30 23:39:25 [__init__.py:662] Selected DeepGemmFp8BlockScaledMMKernel for Fp8LinearMethod
41
+ (Worker_TP0 pid=461) INFO 08-30 23:39:25 [deep_gemm.py:191] deep_gemm not found in site-packages, trying vendored vllm.third_party.deep_gemm
42
+ (Worker_TP0 pid=461) INFO 08-30 23:39:25 [deep_gemm.py:218] DeepGEMM PDL enabled on vllm.third_party.deep_gemm.
43
+ (Worker_TP0 pid=461) INFO 08-30 23:39:25 [deep_gemm.py:136] DeepGEMM E8M0 enabled on current platform.
44
+ (Worker_TP0 pid=461) INFO 08-30 23:39:25 [cuda.py:595] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention
45
+ (Worker_TP0 pid=461) INFO 08-30 23:39:25 [mm_encoder_attention.py:375] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention.
46
+ (Worker_TP1 pid=612) INFO 08-30 23:39:26 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
47
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
48
+ (Worker_TP0 pid=461) WARNING 08-30 23:39:26 [vllm.py:1849] max_num_scheduled_tokens is set to 1024 based on the speculative decoding settings. This may lead to suboptimal performance. Consider increasing max_num_batched_tokens to accommodate the additional draft token slots, or decrease num_speculative_tokens.
49
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
50
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [__init__.py:662] Selected TritonFp8BlockScaledMMKernel for Fp8LinearMethod
51
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [cuda.py:536] Using FLASHINFER_MLA_SPARSE_SM120 attention backend out of potential backends: ['FLASHINFER_MLA_SPARSE_SM120'].
52
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [mla_attention.py:475] Using fp8_ds_mla KV cache format for FLASHINFER_MLA_SPARSE_SM120 backend.
53
+ (Worker_TP0 pid=461) WARNING 08-30 23:39:26 [mla_attention.py:551] Sparse MLA impl has no dense-MHA prefill path; using the top-k MQA path only.
54
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [nvfp4.py:291] Using 'FLASHINFER_CUTLASS' NvFp4 MoE backend out of potential backends: ['FLASHINFER_TRTLLM', 'FLASHINFER_CUTEDSL', 'FLASHINFER_CUTLASS', 'VLLM_CUTLASS', 'MARLIN', 'HUMMING', 'EMULATION'].
55
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [__init__.py:662] Selected DeepGemmFp8BlockScaledMMKernel for _Fp8BlockLMHeadMethod
56
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [weight_utils.py:858] Filesystem type for checkpoints: EXT4. Checkpoint size: 173.36 GiB. Available RAM: 396.63 GiB.
57
+ (Worker_TP0 pid=461) INFO 08-30 23:39:26 [weight_utils.py:881] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
58
+ (Worker_TP0 pid=461)
59
+ (Worker_TP0 pid=461)
60
+ (Worker_TP0 pid=461)
61
+ (Worker_TP0 pid=461)
62
+ (Worker_TP0 pid=461)
63
+ (Worker_TP0 pid=461)
64
+ (Worker_TP0 pid=461)
65
+ (Worker_TP0 pid=461)
66
+ (Worker_TP0 pid=461)
67
+ (Worker_TP0 pid=461)
68
+ (Worker_TP0 pid=461)
69
+ (Worker_TP0 pid=461)
70
+ (Worker_TP0 pid=461)
71
+ (Worker_TP0 pid=461)
72
+ (Worker_TP0 pid=461)
73
+ (Worker_TP0 pid=461)
74
+ (Worker_TP0 pid=461)
75
+ (Worker_TP0 pid=461)
76
+ (Worker_TP0 pid=461)
77
+ (Worker_TP0 pid=461)
78
+ (Worker_TP0 pid=461)
79
+ (Worker_TP0 pid=461)
80
+ (Worker_TP0 pid=461)
81
+ (Worker_TP0 pid=461)
82
+ (Worker_TP0 pid=461)
83
+ (Worker_TP0 pid=461)
84
+ (Worker_TP0 pid=461)
85
+ (Worker_TP0 pid=461)
86
+ (Worker_TP0 pid=461)
87
+ (Worker_TP0 pid=461)
88
+ (Worker_TP0 pid=461)
89
+ (Worker_TP0 pid=461)
90
+ (Worker_TP0 pid=461)
91
+ (Worker_TP0 pid=461)
92
+ (Worker_TP0 pid=461)
93
+ (Worker_TP0 pid=461)
94
+ (Worker_TP0 pid=461)
95
+ (Worker_TP0 pid=461)
96
+ (Worker_TP0 pid=461)
97
+ (Worker_TP0 pid=461)
98
+ (Worker_TP0 pid=461)
99
+ (Worker_TP0 pid=461)
100
+ (Worker_TP0 pid=461)
101
+ (Worker_TP0 pid=461)
102
+ (Worker_TP0 pid=461)
103
+ (Worker_TP0 pid=461)
104
+ (Worker_TP0 pid=461)
105
+ (Worker_TP0 pid=461)
106
+ (Worker_TP0 pid=461)
107
+ (Worker_TP0 pid=461)
108
+ (Worker_TP0 pid=461)
109
+ (Worker_TP0 pid=461)
110
+ (Worker_TP0 pid=461)
111
+ (Worker_TP0 pid=461)
112
+ (Worker_TP0 pid=461)
113
+ (Worker_TP0 pid=461)
114
+ (Worker_TP0 pid=461)
115
+ (Worker_TP0 pid=461)
116
+ (Worker_TP0 pid=461)
117
+ (Worker_TP0 pid=461)
118
+ (Worker_TP0 pid=461)
119
+ (Worker_TP0 pid=461)
120
+ (Worker_TP0 pid=461)
121
+ (Worker_TP0 pid=461)
122
+ (Worker_TP0 pid=461)
123
+ (Worker_TP0 pid=461)
124
+ (Worker_TP0 pid=461)
125
+ (Worker_TP0 pid=461)
126
+ (Worker_TP0 pid=461)
127
+ (Worker_TP0 pid=461)
128
+ (Worker_TP0 pid=461)
129
+ (Worker_TP0 pid=461)
130
+ (Worker_TP0 pid=461)
131
+ (Worker_TP0 pid=461)
132
+ (Worker_TP0 pid=461)
133
+ (Worker_TP0 pid=461)
134
+ (Worker_TP0 pid=461)
135
+ (Worker_TP0 pid=461)
136
+ (Worker_TP0 pid=461)
137
+ (Worker_TP0 pid=461)
138
+ (Worker_TP0 pid=461)
139
+ (Worker_TP0 pid=461)
140
+ (Worker_TP0 pid=461)
141
+ (Worker_TP0 pid=461)
142
+ (Worker_TP0 pid=461)
143
+ (Worker_TP0 pid=461)
144
+ (Worker_TP0 pid=461)
145
+ (Worker_TP0 pid=461)
146
+ (Worker_TP0 pid=461)
147
+ (Worker_TP0 pid=461)
148
+ (Worker_TP0 pid=461)
149
+ (Worker_TP0 pid=461)
150
+ (Worker_TP0 pid=461)
151
+ (Worker_TP0 pid=461)
152
+ (Worker_TP0 pid=461)
153
+ (Worker_TP0 pid=461)
154
+ (Worker_TP0 pid=461)
155
+ (Worker_TP0 pid=461)
156
+ (Worker_TP0 pid=461)
157
+ (Worker_TP0 pid=461)
158
+ (Worker_TP0 pid=461)
159
+ (Worker_TP0 pid=461)
160
+ (Worker_TP0 pid=461)
161
+ (Worker_TP0 pid=461)
162
+ (Worker_TP0 pid=461)
163
+ (Worker_TP0 pid=461)
164
+ (Worker_TP0 pid=461)
165
+ (Worker_TP0 pid=461)
166
+ (Worker_TP0 pid=461)
167
+ (Worker_TP0 pid=461)
168
+ (Worker_TP0 pid=461)
169
+ (Worker_TP0 pid=461)
170
+ (Worker_TP0 pid=461)
171
+ (Worker_TP0 pid=461)
172
+ (Worker_TP0 pid=461)
173
+ (Worker_TP0 pid=461)
174
+ (Worker_TP0 pid=461)
175
+ (Worker_TP0 pid=461)
176
+ (Worker_TP0 pid=461)
177
+ (Worker_TP0 pid=461)
178
+ (Worker_TP0 pid=461)
179
+ (Worker_TP0 pid=461) INFO 08-30 23:40:00 [default_loader.py:430] Loading weights took 33.52 seconds
180
+ [rank1]:[W830 23:40:02.195190924 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 635437056 bytes (free: 14155776, total: 101975851008).
181
+ (Worker_TP0 pid=461) WARNING 08-30 23:40:02 [modelopt.py:1534] w1_weight_scale_2 must match w3_weight_scale_2. Accuracy may be affected.
182
+ (Worker_TP0 pid=461) INFO 08-30 23:40:02 [nvfp4.py:564] Using MoEPrepareAndFinalizeNoDPEPModular
183
+ [rank0]:[W830 23:40:02.661052083 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 635437056 bytes (free: 14155776, total: 101975851008).
184
+ (Worker_TP0 pid=461) INFO 08-30 23:40:02 [weight_utils.py:858] Filesystem type for checkpoints: EXT4. Checkpoint size: 173.36 GiB. Available RAM: 396.49 GiB.
185
+ (Worker_TP0 pid=461)
186
+ (Worker_TP0 pid=461)
187
+ (Worker_TP0 pid=461)
188
+ (Worker_TP0 pid=461)
189
+ (Worker_TP0 pid=461)
190
+ (Worker_TP0 pid=461)
191
+ (Worker_TP0 pid=461)
192
+ (Worker_TP0 pid=461)
193
+ (Worker_TP0 pid=461)
194
+ (Worker_TP0 pid=461)
195
+ (Worker_TP0 pid=461)
196
+ (Worker_TP0 pid=461)
197
+ (Worker_TP0 pid=461)
198
+ (Worker_TP0 pid=461)
199
+ (Worker_TP0 pid=461)
200
+ (Worker_TP0 pid=461)
201
+ (Worker_TP0 pid=461)
202
+ (Worker_TP0 pid=461)
203
+ (Worker_TP0 pid=461)
204
+ (Worker_TP0 pid=461) INFO 08-30 23:40:05 [default_loader.py:430] Loading weights took 2.58 seconds
205
+ (Worker_TP0 pid=461) WARNING 08-30 23:40:05 [speculator.py:191] Draft model Glm5NextMTP does not support external multimodal embeddings. Embeddings from the target model will not be passed to the drafter; using text-only draft inputs instead.
206
+ (Worker_TP1 pid=612) INFO 08-30 23:40:05 [model_runner.py:374] Model loading took 87.28 GiB and 40.227077 seconds
207
+ (Worker_TP1 pid=612) INFO 08-30 23:40:05 [interface.py:635] Setting kv cache block size to 64 for DEEPSEEK_V32_INDEXER backend.
208
+ (Worker_TP1 pid=612) INFO 08-30 23:40:05 [interface.py:926] Setting attention block size to 5120 tokens to ensure that attention page size is >= mamba page size.
209
+ (Worker_TP1 pid=612) INFO 08-30 23:40:05 [interface.py:950] Padding mamba page size by 3.54% to ensure that mamba page size and attention page size are exactly equal.
210
+ (Worker_TP0 pid=461) INFO 08-30 23:40:06 [model_runner.py:374] Model loading took 87.28 GiB and 40.743755 seconds
211
+ (Worker_TP0 pid=461) INFO 08-30 23:40:06 [topk_topp_sampler.py:62] Using FlashInfer for top-p & top-k sampling.
212
+ (Worker_TP0 pid=461) INFO 08-30 23:40:06 [interface.py:635] Setting kv cache block size to 64 for DEEPSEEK_V32_INDEXER backend.
213
+ (Worker_TP0 pid=461) INFO 08-30 23:40:06 [interface.py:926] Setting attention block size to 5120 tokens to ensure that attention page size is >= mamba page size.
214
+ (Worker_TP0 pid=461) INFO 08-30 23:40:06 [interface.py:950] Padding mamba page size by 3.54% to ensure that mamba page size and attention page size are exactly equal.
215
+ (EngineCore pid=327) INFO 08-30 23:40:06 [torch_utils.py:262] Reducing Torch threads from 24 to 1 for serving. Set OMP_NUM_THREADS in the external environment to override.
216
+ (Worker_TP0 pid=461) INFO 08-30 23:40:07 [encoder_runner.py:119] Encoder cache will be initialized with a budget of 7921 tokens, and profiled with 1 image items of the maximum feature size.
217
+ (Worker_TP1 pid=612) 2026-08-30 23:40:16 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
218
+ (Worker_TP1 pid=612) 2026-08-30 23:40:20 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
219
+ (Worker_TP1 pid=612) WARNING 08-30 23:40:21 [rocm.py:42] Failed to import from amdsmi with ModuleNotFoundError("No module named 'amdsmi'")
220
+ (Worker_TP1 pid=612) WARNING 08-30 23:40:21 [rocm.py:47] Failed to import from vllm._C with ModuleNotFoundError("No module named 'vllm._C'")
221
+ (Worker_TP1 pid=612) WARNING 08-30 23:40:21 [rocm.py:58] Failed to import from vllm._rocm_C with ModuleNotFoundError("No module named 'vllm._rocm_C'")
222
+ (Worker_TP1 pid=612) INFO 08-30 23:40:21 [fp8_utils.py:843] Using configuration from /usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/quantization/utils/configs/N=12576,K=4096,device_name=NVIDIA_RTX_PRO_6000_Blackwell_Workstation_Edition,dtype=fp8_w8a8,block_shape=[32,32].json for W8A8 Block FP8 kernel.
223
+ (Worker_TP0 pid=461) 2026-08-30 23:40:16 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
224
+ (Worker_TP0 pid=461) 2026-08-30 23:40:20 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
225
+ (Worker_TP0 pid=461) WARNING 08-30 23:40:21 [rocm.py:42] Failed to import from amdsmi with ModuleNotFoundError("No module named 'amdsmi'")
226
+ (Worker_TP0 pid=461) WARNING 08-30 23:40:21 [rocm.py:47] Failed to import from vllm._C with ModuleNotFoundError("No module named 'vllm._C'")
227
+ (Worker_TP0 pid=461) WARNING 08-30 23:40:21 [rocm.py:58] Failed to import from vllm._rocm_C with ModuleNotFoundError("No module named 'vllm._rocm_C'")
228
+ (Worker_TP0 pid=461) INFO 08-30 23:40:21 [fp8_utils.py:843] Using configuration from /usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/quantization/utils/configs/N=12576,K=4096,device_name=NVIDIA_RTX_PRO_6000_Blackwell_Workstation_Edition,dtype=fp8_w8a8,block_shape=[32,32].json for W8A8 Block FP8 kernel.
229
+ (Worker_TP1 pid=612) /usr/local/lib/python3.12/dist-packages/triton/language/core.py:2284: UserWarning: tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.
230
+ (Worker_TP1 pid=612) warn("tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.")
231
+ (Worker_TP0 pid=461) /usr/local/lib/python3.12/dist-packages/triton/language/core.py:2284: UserWarning: tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.
232
+ (Worker_TP0 pid=461) warn("tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.")
233
+ [23:40:21] : Warning: T.vectorized loop over `i_hci` with extent 4 is lowered as a serial loop because TileLang could not find a valid vectorization plan. Scalar accumulator updates inside the loop are a common cause; move reductions to T.unroll or T.serial if this is intended.
234
+ [23:40:21] : Warning: T.vectorized loop over `i_hci` with extent 4 is lowered as a serial loop because TileLang could not find a valid vectorization plan. Scalar accumulator updates inside the loop are a common cause; move reductions to T.unroll or T.serial if this is intended.
235
+ (Worker_TP0 pid=461) 2026-08-30 23:40:21 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_post_tilelang` with `out_idx=None`
236
+ (Worker_TP0 pid=461) 2026-08-30 23:40:22 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_post_tilelang`
237
+ (Worker_TP0 pid=461) INFO 08-30 23:40:25 [gpu_worker.py:492] Initial free memory 93.92 GiB, reserved 3.24 GiB memory for KV Cache as specified by kv_cache_memory_bytes config and skipped memory profiling. This does not respect the gpu_memory_utilization config. Only use kv_cache_memory_bytes config when you want manual control of KV cache memory size. If OOM'ed, check the difference of initial free memory between the current run and the previous run where kv_cache_memory_bytes is suggested and update it correspondingly.
238
+ (Worker_TP1 pid=612) 2026-08-30 23:40:21 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_post_tilelang` with `out_idx=None`
239
+ (Worker_TP1 pid=612) 2026-08-30 23:40:22 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_post_tilelang`
240
+ (Worker_TP1 pid=612) INFO 08-30 23:40:25 [gpu_worker.py:492] Initial free memory 93.92 GiB, reserved 3.24 GiB memory for KV Cache as specified by kv_cache_memory_bytes config and skipped memory profiling. This does not respect the gpu_memory_utilization config. Only use kv_cache_memory_bytes config when you want manual control of KV cache memory size. If OOM'ed, check the difference of initial free memory between the current run and the previous run where kv_cache_memory_bytes is suggested and update it correspondingly.
241
+ (EngineCore pid=327) INFO 08-30 23:40:25 [kv_cache_utils.py:2141] GPU KV cache size: 547,486 tokens, Maximum concurrency for 524,288 tokens per request: 1.04x
242
+ (Worker_TP0 pid=461) INFO 08-30 23:40:25 [indexer.py:762] DSA indexer decode path: use_flattening=False supports_varlen=False (next_n=2, use_fp4_indexer_cache=False)
243
+ (Worker_TP0 pid=461) INFO 08-30 23:40:25 [kernel_warmup.py:95] Warming up ll_bf16 router GEMM kernels for shapes: ((4096, 288),).
244
+ (Worker_TP0 pid=461)
245
+ (Worker_TP0 pid=461) INFO 08-30 23:40:26 [kernel_warmup.py:170] Skipping FlashInfer autotune because it is disabled.
246
+ (Worker_TP0 pid=461) INFO 08-30 23:40:26 [breakable_cudagraph.py:288] Breakable CUDA graph enabled
247
+ (Worker_TP0 pid=461)
248
 
249
 
250
+ (Worker_TP1 pid=612) warn("tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.")
251
+ /usr/local/lib/python3.12/dist-packages/triton/language/core.py:2284: UserWarning: tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.
252
+ (Worker_TP0 pid=461) warn("tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.")
253
+ (Worker_TP0 pid=461)
254
 
255
 
256
 
257
 
258
+ (Worker_TP0 pid=461)
259
+ (Worker_TP0 pid=461) 2026-08-30 23:40:31 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
260
+ (Worker_TP0 pid=461) 2026-08-30 23:41:13 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_fused_tilelang` with `out_idx=None`
261
+ (Worker_TP0 pid=461) 2026-08-30 23:41:13 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_fused_tilelang`
262
+ (Worker_TP0 pid=461) 2026-08-30 23:41:14 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
263
+ (Worker_TP0 pid=461) 2026-08-30 23:41:18 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
264
+
265
+ (Worker_TP1 pid=612) 2026-08-30 23:40:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
266
+ (Worker_TP1 pid=612) 2026-08-30 23:40:31 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
267
+ (Worker_TP1 pid=612) 2026-08-30 23:41:13 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_fused_tilelang` with `out_idx=None`
268
+ (Worker_TP1 pid=612) 2026-08-30 23:41:13 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_fused_tilelang`
269
+ (Worker_TP1 pid=612) 2026-08-30 23:41:14 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
270
+ (Worker_TP1 pid=612) 2026-08-30 23:41:18 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
271
+ (Worker_TP1 pid=612) INFO 08-30 23:41:22 [speculator.py:148] Capturing model for speculator...
272
+ (Worker_TP0 pid=461) INFO 08-30 23:41:22 [speculator.py:148] Capturing model for speculator...
273
+ (Worker_TP0 pid=461)
274
+ (Worker_TP0 pid=461)
275
+ (Worker_TP1 pid=612) INFO 08-30 23:41:23 [model_runner.py:1012] Graph capturing finished in 57 secs, took 0.61 GiB
276
+ (Worker_TP0 pid=461) INFO 08-30 23:41:23 [model_runner.py:1012] Graph capturing finished in 57 secs, took 0.61 GiB
277
+ (Worker_TP0 pid=461) INFO 08-30 23:41:26 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
278
+ (Worker_TP1 pid=612) INFO 08-30 23:41:26 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
279
+ (EngineCore pid=327) INFO 08-30 23:41:26 [shm_broadcast.py:801] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization).
280
+ (Worker_TP0 pid=461) INFO 08-30 23:41:26 [torch_utils.py:262] Reducing Torch threads from 24 to 1 for serving. Set OMP_NUM_THREADS in the external environment to override.
281
+ (EngineCore pid=327) INFO 08-30 23:41:26 [core.py:363] init engine (profile, create kv cache, warmup model) took 80.44 s
282
+ (EngineCore pid=327) INFO 08-30 23:41:30 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
283
+ (EngineCore pid=327) WARNING 08-30 23:41:30 [vllm.py:1849] max_num_scheduled_tokens is set to 1024 based on the speculative decoding settings. This may lead to suboptimal performance. Consider increasing max_num_batched_tokens to accommodate the additional draft token slots, or decrease num_speculative_tokens.
284
+ (EngineCore pid=327) INFO 08-30 23:41:30 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
285
+ (APIServer pid=59) INFO 08-30 23:41:31 [api_server.py:678] Supported tasks: ['generate']
286
+ (APIServer pid=59) INFO 08-30 23:41:31 [parser_manager.py:37] "auto" tool choice has been enabled.
287
+ (APIServer pid=59) INFO 08-30 23:41:32 [hf.py:551] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this.
288
+ (APIServer pid=59) INFO 08-30 23:41:34 [base.py:235] Multi-modal warmup completed in 1.661s
289
+ (APIServer pid=59) INFO 08-30 23:41:36 [base.py:235] Readonly multi-modal warmup completed in 1.725s
290
+ (APIServer pid=59) WARNING 08-30 23:41:36 [model.py:1720] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 1.0, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
291
+ (APIServer pid=59) INFO 08-30 23:41:39 [api_server.py:682] Starting vLLM server on http://0.0.0.0:18998
292
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:51] Available routes are:
293
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /openapi.json, Methods: HEAD, GET
294
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /docs, Methods: HEAD, GET
295
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /docs/oauth2-redirect, Methods: HEAD, GET
296
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /redoc, Methods: HEAD, GET
297
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /load, Methods: GET
298
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /version, Methods: GET
299
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /health, Methods: GET
300
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /metrics, Methods: GET
301
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /tokenize, Methods: POST
302
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /detokenize, Methods: POST
303
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/models, Methods: GET
304
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /ping, Methods: GET
305
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /ping, Methods: POST
306
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /invocations, Methods: POST
307
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/chat/completions, Methods: POST
308
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/chat/completions/batch, Methods: POST
309
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/responses, Methods: POST
310
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/responses/{response_id}, Methods: GET
311
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/responses/{response_id}/cancel, Methods: POST
312
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/completions, Methods: POST
313
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/messages, Methods: POST
314
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/messages/count_tokens, Methods: POST
315
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /generative_scoring, Methods: POST
316
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /scale_elastic_ep, Methods: POST
317
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /is_scaling_elastic_ep, Methods: POST
318
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/chat/completions/render, Methods: POST
319
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/completions/render, Methods: POST
320
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/chat/completions/derender, Methods: POST
321
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/completions/derender, Methods: POST
322
+ (APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /inference/v1/generate, Methods: POST
323
+ (APIServer pid=59) INFO: Started server process [59]
324
+ (APIServer pid=59) INFO: Waiting for application startup.
325
+ (APIServer pid=59) INFO: Application startup complete.
326
+ (APIServer pid=59) INFO: 127.0.0.1:56012 - "GET /health HTTP/1.1" 200 OK
327
+ (Worker_TP1 pid=612) INFO 08-30 23:40:05 [model_runner.py:374] Model loading took 87.28 GiB and 40.227077 seconds
328
+ (Worker_TP0 pid=461) INFO 08-30 23:40:06 [model_runner.py:374] Model loading took 87.28 GiB and 40.743755 seconds
329
+ (EngineCore pid=327) INFO 08-30 23:40:25 [kv_cache_utils.py:2141] GPU KV cache size: 547,486 tokens, Maximum concurrency for 524,288 tokens per request: 1.04x
330
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:41 [jit_monitor.py:135] Triton kernel JIT compilation during inference: BuildPrefillChunkMetadataKernel.kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
331
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:41 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _w8a8_triton_block_scaled_mm. This causes a latency spike; consider extending warmup to cover this shape/config.
332
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:42 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _kpool_softmax_rotate_write_cache_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
333
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:42 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _compute_local_logits_stats_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
334
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:42 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _rejection_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
335
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:42 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _resample_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
336
+ (APIServer pid=59) INFO: 127.0.0.1:56028 - "POST /v1/chat/completions HTTP/1.1" 200 OK
337
+ TEXT: PASS (391)
338
+ (APIServer pid=59) INFO: 127.0.0.1:56040 - "POST /v1/chat/completions HTTP/1.1" 200 OK
339
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:43 [jit_monitor.py:135] Triton kernel JIT compilation during inference: rotary_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
340
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:43 [jit_monitor.py:135] TileLang JIT compilation during inference: mhc_pre_big_fuse_with_norm_tilelang. This causes a latency spike; consider extending warmup to cover this shape/config.
341
+ (APIServer pid=59) INFO: 127.0.0.1:56054 - "POST /v1/chat/completions HTTP/1.1" 200 OK
342
+ (APIServer pid=59) INFO: 127.0.0.1:56070 - "POST /v1/chat/completions HTTP/1.1" 200 OK
343
+ (APIServer pid=59) INFO: 127.0.0.1:56074 - "POST /v1/chat/completions HTTP/1.1" 200 OK
344
+ (APIServer pid=59) INFO 08-30 23:41:49 [loggers.py:310] Engine 000: Avg prompt throughput: 34.4 tokens/s, Avg generation throughput: 6.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 6.8%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
345
+ (APIServer pid=59) INFO 08-30 23:41:49 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.96, Accepted throughput: 2.94 tokens/s, Drafted throughput: 3.05 tokens/s, Accepted: 55 tokens, Drafted: 57 tokens, Per-position acceptance rate: 0.965, Avg Draft acceptance rate: 96.5%
346
+ (APIServer pid=59) INFO: 127.0.0.1:56078 - "POST /v1/chat/completions HTTP/1.1" 200 OK
347
+ (APIServer pid=59) INFO: 127.0.0.1:56090 - "POST /v1/chat/completions HTTP/1.1" 200 OK
348
+ (APIServer pid=59) INFO: 127.0.0.1:56096 - "POST /v1/chat/completions HTTP/1.1" 200 OK
349
+ (APIServer pid=59) INFO: 127.0.0.1:56108 - "POST /v1/chat/completions HTTP/1.1" 200 OK
350
+ (APIServer pid=59) INFO: 127.0.0.1:56118 - "POST /v1/chat/completions HTTP/1.1" 200 OK
351
+ (APIServer pid=59) INFO: 127.0.0.1:56124 - "POST /v1/chat/completions HTTP/1.1" 200 OK
352
+ (APIServer pid=59) INFO: 127.0.0.1:51220 - "POST /v1/chat/completions HTTP/1.1" 200 OK
353
+ (Worker_TP0 pid=461) 2026-08-30 23:41:43 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
354
+ (Worker_TP0 pid=461) 2026-08-30 23:41:48 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
355
+ (Worker_TP0 pid=461) WARNING 08-30 23:41:51 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _kpool_tail_seed_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
356
+ (APIServer pid=59) INFO: 127.0.0.1:51222 - "POST /v1/chat/completions HTTP/1.1" 200 OK
357
+ (APIServer pid=59) INFO: 127.0.0.1:51238 - "POST /v1/chat/completions HTTP/1.1" 200 OK
358
+ (APIServer pid=59) INFO: 127.0.0.1:51242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
359
+ (APIServer pid=59) INFO: 127.0.0.1:51244 - "POST /v1/chat/completions HTTP/1.1" 200 OK
360
+ (APIServer pid=59) INFO: 127.0.0.1:51256 - "POST /v1/chat/completions HTTP/1.1" 200 OK
361
+ (APIServer pid=59) INFO: 127.0.0.1:51266 - "POST /v1/chat/completions HTTP/1.1" 200 OK
362
+ (APIServer pid=59) INFO: 127.0.0.1:51272 - "POST /v1/chat/completions HTTP/1.1" 200 OK
363
+ (APIServer pid=59) INFO: 127.0.0.1:51278 - "POST /v1/chat/completions HTTP/1.1" 200 OK
364
+ (APIServer pid=59) INFO: 127.0.0.1:51284 - "POST /v1/chat/completions HTTP/1.1" 200 OK
365
+ [img01.png] OK img_toks=258 ans='Red'
366
+ [img02.png] OK img_toks=258 ans='SATURN'
367
+ [img03.png] OK img_toks=258 ans='4'
368
+ [img04.png] OK img_toks=258 ans='Triangle'
369
+ [img05.png] OK img_toks=258 ans='7318'
370
+ [img06.png] OK img_toks=258 ans='3'
371
+ [img07.png] OK img_toks=258 ans='Circle'
372
+ [img08.png] OK img_toks=258 ans='MANGO'
373
+ [img09.png] OK img_toks=258 ans='5'
374
+ [img10.png] OK img_toks=258 ans='The black square is on the **left** side of the image.'
375
+ VISION VERDICT: PASS (10/10 correct, gate >=10, hard_fail=False)
376
+ wrote results-vision-512kconf-8666.json
377
+ === 1x ~501k fresh prefill ===
378
+ (APIServer pid=59) INFO: 127.0.0.1:51296 - "POST /tokenize HTTP/1.1" 200 OK
379
+ (APIServer pid=59) INFO: 127.0.0.1:51304 - "POST /tokenize HTTP/1.1" 200 OK
380
+ (APIServer pid=59) INFO: 127.0.0.1:51318 - "POST /tokenize HTTP/1.1" 200 OK
381
+ (APIServer pid=59) INFO: 127.0.0.1:51332 - "POST /tokenize HTTP/1.1" 200 OK
382
+ (APIServer pid=59) INFO: 127.0.0.1:51340 - "POST /tokenize HTTP/1.1" 200 OK
383
+ (APIServer pid=59) INFO 08-30 23:41:59 [loggers.py:310] Engine 000: Avg prompt throughput: 249.9 tokens/s, Avg generation throughput: 61.3 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
384
+ (APIServer pid=59) INFO 08-30 23:41:59 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.89, Accepted throughput: 28.20 tokens/s, Drafted throughput: 31.79 tokens/s, Accepted: 283 tokens, Drafted: 319 tokens, Per-position acceptance rate: 0.887, Avg Draft acceptance rate: 88.7%
385
+ (APIServer pid=59) INFO: 127.0.0.1:51346 - "POST /tokenize HTTP/1.1" 200 OK
386
+ (APIServer pid=59) INFO: 127.0.0.1:51354 - "POST /tokenize HTTP/1.1" 200 OK
387
+ (APIServer pid=59) INFO: 127.0.0.1:51358 - "POST /tokenize HTTP/1.1" 200 OK
388
+ (APIServer pid=59) INFO: 127.0.0.1:51366 - "POST /tokenize HTTP/1.1" 200 OK
389
+ (APIServer pid=59) INFO: 127.0.0.1:46572 - "POST /tokenize HTTP/1.1" 200 OK
390
+ (APIServer pid=59) INFO: 127.0.0.1:46580 - "POST /tokenize HTTP/1.1" 200 OK
391
+ (APIServer pid=59) INFO: 127.0.0.1:46582 - "POST /tokenize HTTP/1.1" 200 OK
392
+ (APIServer pid=59) INFO: 127.0.0.1:46592 - "POST /tokenize HTTP/1.1" 200 OK
393
+ (APIServer pid=59) INFO: 127.0.0.1:46608 - "POST /tokenize HTTP/1.1" 200 OK
394
+ (APIServer pid=59) INFO: 127.0.0.1:46616 - "POST /tokenize HTTP/1.1" 200 OK
395
+ (APIServer pid=59) INFO 08-30 23:42:09 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 8.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
396
+ [rank1]:[W830 23:42:34.181939325 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 178257920 bytes (free: 114819072, total: 101975851008).
397
+ [rank0]:[W830 23:42:34.182028155 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 178257920 bytes (free: 127401984, total: 101975851008).
398
+ [rank1]:[W830 23:42:38.152985005 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 203423744 bytes (free: 158859264, total: 101975851008).
399
+ [rank0]:[W830 23:42:38.490482835 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 205520896 bytes (free: 68681728, total: 101975851008).
400
+ [rank1]:[W830 23:42:41.840248434 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 226492416 bytes (free: 217579520, total: 101975851008).
401
+ [rank0]:[W830 23:42:42.181948344 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 228589568 bytes (free: 211288064, total: 101975851008).
402
+ [rank1]:[W830 23:42:45.224976732 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 247463936 bytes (free: 211288064, total: 101975851008).
403
+ [rank0]:[W830 23:42:45.571221181 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 249561088 bytes (free: 207093760, total: 101975851008).
404
+ [rank1]:[W830 23:42:48.301529558 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 266338304 bytes (free: 267911168, total: 101975851008).
405
+ [rank0]:[W830 23:42:48.649412232 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 268435456 bytes (free: 265814016, total: 101975851008).
406
+ [rank1]:[W830 23:42:51.403194296 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 285212672 bytes (free: 98041856, total: 101975851008).
407
+ [rank0]:[W830 23:42:51.753405922 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 287309824 bytes (free: 95944704, total: 101975851008).
408
+ [rank1]:[W830 23:42:54.183804863 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 301989888 bytes (free: 230162432, total: 101975851008).
409
+ [rank0]:[W830 23:42:54.536068703 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 304087040 bytes (free: 230162432, total: 101975851008).
410
+ [rank1]:[W830 23:42:56.986131102 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 318767104 bytes (free: 95944704, total: 101975851008).
411
+ [rank0]:[W830 23:42:57.340623246 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 320864256 bytes (free: 95944704, total: 101975851008).
412
+ [rank1]:[W830 23:42:59.449005812 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 333447168 bytes (free: 295174144, total: 101975851008).
413
+ [rank0]:[W830 23:42:59.805219431 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 335544320 bytes (free: 297271296, total: 101975851008).
414
+ [rank1]:[W830 23:43:01.931529503 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 348127232 bytes (free: 192413696, total: 101975851008).
415
+ [rank0]:[W830 23:43:02.290645974 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 350224384 bytes (free: 194510848, total: 101975851008).
416
+ [rank1]:[W830 23:43:04.427191813 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 362807296 bytes (free: 89653248, total: 101975851008).
417
+ [rank0]:[W830 23:43:04.788776567 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 364904448 bytes (free: 91750400, total: 101975851008).
418
+ [rank1]:[W830 23:43:06.582338268 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 375390208 bytes (free: 362283008, total: 101975851008).
419
+ [rank0]:[W830 23:43:06.945909107 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 377487360 bytes (free: 366477312, total: 101975851008).
420
+ [rank1]:[W830 23:43:08.749475760 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 387973120 bytes (free: 286785536, total: 101975851008).
421
+ [rank0]:[W830 23:43:09.114796003 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 390070272 bytes (free: 290979840, total: 101975851008).
422
+ [rank1]:[W830 23:43:10.928961237 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 400556032 bytes (free: 211288064, total: 101975851008).
423
+ [rank0]:[W830 23:43:11.298048332 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 402653184 bytes (free: 215482368, total: 101975851008).
424
+ [rank1]:[W830 23:43:13.120925840 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 413138944 bytes (free: 135790592, total: 101975851008).
425
+ [rank0]:[W830 23:43:13.491285390 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 415236096 bytes (free: 139984896, total: 101975851008).
426
+ [rank1]:[W830 23:43:15.324590440 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 425721856 bytes (free: 60293120, total: 101975851008).
427
+ [rank0]:[W830 23:43:15.695674317 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 427819008 bytes (free: 64487424, total: 101975851008).
428
+ [rank1]:[W830 23:43:17.172276925 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 436207616 bytes (free: 421003264, total: 101975851008).
429
+ [rank0]:[W830 23:43:17.545006047 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 438304768 bytes (free: 427294720, total: 101975851008).
430
+ [rank1]:[W830 23:43:19.028108396 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 446693376 bytes (free: 368574464, total: 101975851008).
431
+ [rank0]:[W830 23:43:19.401915915 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 448790528 bytes (free: 374865920, total: 101975851008).
432
+ [rank1]:[W830 23:43:20.896195487 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 457179136 bytes (free: 316145664, total: 101975851008).
433
+ [rank0]:[W830 23:43:21.270943542 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 459276288 bytes (free: 322437120, total: 101975851008).
434
+ [rank1]:[W830 23:43:22.740643999 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 467664896 bytes (free: 263716864, total: 101975851008).
435
+ [rank0]:[W830 23:43:23.108960854 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 469762048 bytes (free: 270008320, total: 101975851008).
436
+ [rank1]:[W830 23:43:24.564337548 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 478150656 bytes (free: 211288064, total: 101975851008).
437
+ [rank0]:[W830 23:43:24.932404841 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 480247808 bytes (free: 217579520, total: 101975851008).
438
+ [rank1]:[W830 23:43:26.388998346 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 488636416 bytes (free: 158859264, total: 101975851008).
439
+ [rank0]:[W830 23:43:26.759605073 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 490733568 bytes (free: 165150720, total: 101975851008).
440
+ [rank1]:[W830 23:43:28.217846118 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 499122176 bytes (free: 106430464, total: 101975851008).
441
+ [rank0]:[W830 23:43:28.590325529 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 501219328 bytes (free: 112721920, total: 101975851008).
442
+ [rank1]:[W830 23:43:30.057423902 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 509607936 bytes (free: 54001664, total: 101975851008).
443
+ [rank0]:[W830 23:43:30.428232634 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 511705088 bytes (free: 60293120, total: 101975851008).
444
+ (APIServer pid=59) INFO: 127.0.0.1:46622 - "POST /v1/chat/completions HTTP/1.1" 200 OK
445
+ (APIServer pid=59) INFO: 127.0.0.1:50372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
446
+ (APIServer pid=59) INFO: 127.0.0.1:50386 - "POST /v1/chat/completions HTTP/1.1" 200 OK
447
+ calibration: 0.18182 tok/char on 9471 chars
448
+ [round 1] built prompt: 434531 tokens (2390199 chars)
449
+ [round 1] extended to 500974 tokens
450
+ [round 1] prompt_tokens=501006 completion=55 wall=84.7s finish=stop ok=True
451
+ [round 1] reply: 'ACKNOWLEDGED. The text appears to be random filler content with no meaningful message.'
452
+ [round 1] decode at 501006 ctx: 61 tok in 3.0s (1-tok call 2.7s) -> ~272.1 tok/s
453
+ PASS 1/1
454
+ === post-500k vision re-probe ===
455
+ (APIServer pid=59) INFO: 127.0.0.1:50388 - "POST /v1/chat/completions HTTP/1.1" 200 OK
456
+ (APIServer pid=59) INFO: 127.0.0.1:50398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
457
+ (APIServer pid=59) INFO 08-30 23:43:39 [loggers.py:310] Engine 000: Avg prompt throughput: 52082.9 tokens/s, Avg generation throughput: 40.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 6.8%, Prefix cache hit rate: 65.2%, MM cache hit rate: 0.0%
458
+ (APIServer pid=59) INFO 08-30 23:43:39 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.96, Accepted throughput: 1.94 tokens/s, Drafted throughput: 2.03 tokens/s, Accepted: 194 tokens, Drafted: 203 tokens, Per-position acceptance rate: 0.956, Avg Draft acceptance rate: 95.6%
459
+ (APIServer pid=59) INFO: 127.0.0.1:50408 - "POST /v1/chat/completions HTTP/1.1" 200 OK
460
+ (APIServer pid=59) INFO: 127.0.0.1:50420 - "POST /v1/chat/completions HTTP/1.1" 200 OK
461
+ (APIServer pid=59) INFO: 127.0.0.1:48808 - "POST /v1/chat/completions HTTP/1.1" 200 OK
462
+ (APIServer pid=59) INFO: 127.0.0.1:48818 - "POST /v1/chat/completions HTTP/1.1" 200 OK
463
+ (APIServer pid=59) INFO: 127.0.0.1:48832 - "POST /v1/chat/completions HTTP/1.1" 200 OK
464
+ (APIServer pid=59) INFO: 127.0.0.1:48848 - "POST /v1/chat/completions HTTP/1.1" 200 OK
465
+ (APIServer pid=59) INFO: 127.0.0.1:48850 - "POST /v1/chat/completions HTTP/1.1" 200 OK
466
+ (APIServer pid=59) INFO: 127.0.0.1:48858 - "POST /v1/chat/completions HTTP/1.1" 200 OK
467
+ (APIServer pid=59) INFO: 127.0.0.1:48874 - "POST /v1/chat/completions HTTP/1.1" 200 OK
468
+ (APIServer pid=59) INFO: 127.0.0.1:48882 - "POST /v1/chat/completions HTTP/1.1" 200 OK
469
+ [PASS] shapes_3red_circles: gold='3' got='3' (1.0s)
470
+ [PASS] shapes_5blue_squares: gold='5' got='5' (1.1s)
471
+ [PASS] shapes_2green_circles: gold='2' got='2' (0.8s)
472
+ [PASS] shapes_4yellow_squares: gold='4' got='4' (1.2s)
473
+ [PASS] text_mango: gold='MANGO' got='MANGO' (0.3s)
474
+ [PASS] text_river7: gold='RIVER 7' got='RIVER 7' (0.3s)
475
+ [PASS] text_zebra_code: gold='ZK-42' got='ZK-42' (0.4s)
476
+ [PASS] text_plum: gold='PLUM BASKET' got='PLUM BASKET' (0.4s)
477
+ [PASS] chart_tallest_mar: gold='MAR' got='MAR' (0.4s)
478
+ [PASS] chart_shortest_q2: gold='Q2' got='Q2' (0.4s)
479
+ [PASS] chart_value_feb: gold='75' got='75' (0.4s)
480
+ [PASS] chart_pie_cats: gold='CATS' got='**CATS**' (0.9s)
481
+ SCORE 512kconf-post500k-8666: 12/12
482
+ wrote /home/ian/h3-v2/glm-max/results-vision-512kconf-post500k-8666.json
483
+ === min free watermark (MiB): 110 ===
484
+ VISFP8-512K done
485
+ [rewarm] posting completion to prod glm-5.3-flash via 8080...
486
+ [rewarm] answer: 391 17 × 23 = 17 × 23
487
+
488
+ Let me compute: 17 × 23 = 17 × 20 + 17 × 3 = 340 + 51 = 391
489
+ [rewarm] running: glm-5.3-flash=ready
490
+ [rewarm] PROD WARM OK
491
+ [bash] released
scripts/convert_fp8visual.py ADDED
@@ -0,0 +1,229 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Derive a block-FP8-visual-tower checkpoint from GLM-5.3-Flash-NVFP4-FP8ATTN-recal.
2
+
3
+ Worker VISION (glm-max push). The r2-era "yields NaN image features" warning in
4
+ model.py (~line 1050) is about inheriting the global fp8 quant_config while the
5
+ checkpoint ships BF16 visual weights WITHOUT weight_scale_inv — i.e. a
6
+ loader/format mismatch, not FP8 numerics of the tower. This conversion ships
7
+ the scales, so with the manifest-gated model.py overlay (vision-test image) the
8
+ tower's Linear layers load through vLLM's proven Fp8LinearMethod block path
9
+ (same serialization as the r2 MLP/attention conversion: weight F8_E4M3 [N,K] +
10
+ weight_scale_inv F32 [N/128,K/128], dequant multiplier amax/448).
11
+
12
+ Converted (124 tensors, 1.0156 GiB BF16 -> ~0.512 GiB, all in shard 120):
13
+ model.visual.blocks.{0..23}.attn.qkv.weight [3072,1024]
14
+ model.visual.blocks.{0..23}.attn.proj.weight [1024,1024]
15
+ model.visual.blocks.{0..23}.mlp.{gate,up,down}_proj.weight
16
+ model.visual.merger.{proj,gate_proj,up_proj,down_proj}.weight
17
+ NOT converted: patch_embed (Conv3d), downsample (Conv2d), all biases/norms
18
+ (34.2 MiB total). Every shape and TP2 shard offset divides by 128 (verified:
19
+ qkv q/k/v sections 1024 each -> 512/rank; merger 10240 -> 5120/rank).
20
+
21
+ Config: prune the three visual ignore patterns; add FP8_BLOCK128
22
+ quantized_layers entries under the RUNTIME lookup prefixes
23
+ ("visual.blocks.N.attn.qkv_proj" — note qkv_proj, multimodal.py renames the
24
+ prefix when quant_config is passed) plus "model.visual." spellings and
25
+ per-shard aliases (strategy-4 fused lookups), mirroring FIT's r4 alias
26
+ precedent. Unlisted visual linears resolve to UnquantizedLinearMethod, so
27
+ pruning the ignore patterns is safe for the conv/norm remainder.
28
+
29
+ Usage: python3 convert_fp8visual.py # recal -> -recal-visfp8
30
+ SRC=<dir> DST=<dir> python3 ... # stack on DRAFT's dir
31
+ """
32
+
33
+ import json
34
+ import os
35
+ import re
36
+ import struct
37
+ import time
38
+
39
+ import ml_dtypes
40
+ import numpy as np
41
+
42
+ SRC = os.environ.get("SRC", "/home/ian/models/GLM-5.3-Flash-NVFP4-FP8ATTN-recal")
43
+ DST = os.environ.get(
44
+ "DST", "/home/ian/models/GLM-5.3-Flash-NVFP4-FP8ATTN-recal-visfp8"
45
+ )
46
+ BLOCK = 128
47
+ DEPTH = 24
48
+
49
+ TARGET_RE = re.compile(
50
+ r"^model\.visual\.("
51
+ r"blocks\.\d+\.attn\.(qkv|proj)"
52
+ r"|blocks\.\d+\.mlp\.(gate|up|down)_proj"
53
+ r"|merger\.(proj|gate_proj|up_proj|down_proj)"
54
+ r")\.weight$"
55
+ )
56
+
57
+ VISUAL_IGNORE = {"model.visual.*", "visual.*", "*.visual.*"}
58
+
59
+
60
+ def read_header(path):
61
+ with open(path, "rb") as f:
62
+ n = struct.unpack("<Q", f.read(8))[0]
63
+ return json.loads(f.read(n)), 8 + n
64
+
65
+
66
+ def quant_block(w32, b=BLOCK):
67
+ n, k = w32.shape
68
+ assert n % b == 0 and k % b == 0, (n, k)
69
+ B = w32.reshape(n // b, b, k // b, b)
70
+ amax = np.abs(B).max(axis=(1, 3), keepdims=True)
71
+ s = np.where(amax == 0, np.float32(1), amax / np.float32(448)).astype(np.float32)
72
+ q = np.clip(B / s, -448, 448).astype(ml_dtypes.float8_e4m3fn)
73
+ return (
74
+ np.ascontiguousarray(q.reshape(n, k)),
75
+ np.ascontiguousarray(s.reshape(B.shape[0], B.shape[2])),
76
+ )
77
+
78
+
79
+ def visual_manifest_entries():
80
+ fp8 = {"quant_algo": "FP8_BLOCK128"}
81
+ mods = []
82
+ for i in range(DEPTH):
83
+ mods += [
84
+ # runtime lookup prefix (multimodal.py: ".qkv_proj" iff quant_config)
85
+ f"visual.blocks.{i}.attn.qkv_proj",
86
+ f"visual.blocks.{i}.attn.qkv",
87
+ f"visual.blocks.{i}.attn.proj",
88
+ f"visual.blocks.{i}.mlp.gate_up_proj",
89
+ f"visual.blocks.{i}.mlp.gate_proj",
90
+ f"visual.blocks.{i}.mlp.up_proj",
91
+ f"visual.blocks.{i}.mlp.down_proj",
92
+ ]
93
+ mods += [
94
+ "visual.merger.proj",
95
+ "visual.merger.gate_up_proj",
96
+ "visual.merger.gate_proj",
97
+ "visual.merger.up_proj",
98
+ "visual.merger.down_proj",
99
+ ]
100
+ ql = {}
101
+ for m in mods:
102
+ ql[m] = fp8
103
+ ql["model." + m] = fp8 # in case apply_vllm_mapper is/isn't applied
104
+ return ql
105
+
106
+
107
+ def main():
108
+ os.makedirs(DST, exist_ok=True)
109
+ idx = json.load(open(f"{SRC}/model.safetensors.index.json"))
110
+ wmap = dict(idx["weight_map"])
111
+ shards = sorted(set(wmap.values()))
112
+
113
+ todo = {}
114
+ for name, sh in wmap.items():
115
+ if TARGET_RE.match(name):
116
+ todo.setdefault(sh, set()).add(name)
117
+ n_targets = sum(len(v) for v in todo.values())
118
+ print(f"{n_targets} target tensors in {len(todo)}/{len(shards)} shards")
119
+ assert n_targets == 124, n_targets
120
+
121
+ t0 = time.time()
122
+ linked = rewritten = 0
123
+ size_delta = 0
124
+ for si, sh in enumerate(shards, 1):
125
+ src, dst = f"{SRC}/{sh}", f"{DST}/{sh}"
126
+ if os.path.exists(dst):
127
+ os.unlink(dst)
128
+ if sh not in todo:
129
+ os.link(src, dst)
130
+ linked += 1
131
+ continue
132
+ hdr, base = read_header(src)
133
+ meta = hdr.get("__metadata__")
134
+ conv = todo[sh]
135
+ order = [k for k in hdr if k != "__metadata__"]
136
+ new_hdr, blobs, off = {}, [], 0
137
+ with open(src, "rb") as f:
138
+ for k in order:
139
+ m = hdr[k]
140
+ s, e = m["data_offsets"]
141
+ f.seek(base + s)
142
+ raw = f.read(e - s)
143
+ if k in conv:
144
+ assert m["dtype"] == "BF16", (k, m["dtype"])
145
+ w = (
146
+ np.frombuffer(raw, dtype=ml_dtypes.bfloat16)
147
+ .reshape(m["shape"])
148
+ .astype(np.float32)
149
+ )
150
+ qw, sc = quant_block(w)
151
+ # round-trip check
152
+ deq = qw.astype(np.float32) * np.kron(
153
+ sc, np.ones((BLOCK, BLOCK), np.float32)
154
+ )
155
+ rel = np.linalg.norm(deq - w) / np.linalg.norm(w)
156
+ if rel > 0.05:
157
+ print(f" WARN {k} round-trip rel err {rel:.4%}")
158
+ for nm, arr, dt in (
159
+ (k, qw, "F8_E4M3"),
160
+ (k[: -len(".weight")] + ".weight_scale_inv", sc, "F32"),
161
+ ):
162
+ b = arr.tobytes()
163
+ new_hdr[nm] = {
164
+ "dtype": dt,
165
+ "shape": list(arr.shape),
166
+ "data_offsets": [off, off + len(b)],
167
+ }
168
+ blobs.append(b)
169
+ off += len(b)
170
+ size_delta += len(b)
171
+ size_delta -= len(raw)
172
+ else:
173
+ new_hdr[k] = {
174
+ "dtype": m["dtype"],
175
+ "shape": m["shape"],
176
+ "data_offsets": [off, off + len(raw)],
177
+ }
178
+ blobs.append(raw)
179
+ off += len(raw)
180
+ if meta is not None:
181
+ new_hdr["__metadata__"] = meta
182
+ hb = json.dumps(new_hdr, separators=(",", ":")).encode()
183
+ pad = (-(8 + len(hb))) % 8
184
+ hb += b" " * pad
185
+ tmp = dst + ".tmp"
186
+ with open(tmp, "wb") as f:
187
+ f.write(struct.pack("<Q", len(hb)))
188
+ f.write(hb)
189
+ for b in blobs:
190
+ f.write(b)
191
+ os.replace(tmp, dst)
192
+ rewritten += 1
193
+ print(
194
+ f"[{si}/{len(shards)}] {sh} rewrote {len(conv)} tensors "
195
+ f"({time.time() - t0:.0f}s)",
196
+ flush=True,
197
+ )
198
+
199
+ for sh, names in todo.items():
200
+ for name in names:
201
+ wmap[name[: -len(".weight")] + ".weight_scale_inv"] = sh
202
+ idx["weight_map"] = wmap
203
+ if "metadata" in idx and "total_size" in idx["metadata"]:
204
+ idx["metadata"]["total_size"] += size_delta
205
+ json.dump(idx, open(f"{DST}/model.safetensors.index.json", "w"))
206
+
207
+ cfg = json.load(open(f"{SRC}/config.json"))
208
+ q = cfg["quantization_config"]
209
+ q["ignore"] = [p for p in q["ignore"] if p not in VISUAL_IGNORE]
210
+ q["quantized_layers"] = {**q["quantized_layers"], **visual_manifest_entries()}
211
+ json.dump(cfg, open(f"{DST}/config.json", "w"), indent=1)
212
+
213
+ for fn in os.listdir(SRC):
214
+ if fn.endswith(".safetensors") or fn in (
215
+ "config.json",
216
+ "model.safetensors.index.json",
217
+ ):
218
+ continue
219
+ s, d = f"{SRC}/{fn}", f"{DST}/{fn}"
220
+ if os.path.isfile(s) and not os.path.exists(d):
221
+ os.link(s, d)
222
+ print(
223
+ f"DONE linked={linked} rewritten={rewritten} "
224
+ f"size_delta={size_delta / 2**30:.3f} GiB in {time.time() - t0:.0f}s"
225
+ )
226
+
227
+
228
+ if __name__ == "__main__":
229
+ main()
scripts/results-tf-8666.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"port": 18998, "model": "glm-max", "arith_correct": true, "arith_raw": "391", "long_gen_tokens": 2400, "long_gen_degenerate": false, "long_gen_tail": "with 226 containers... Actually Gateway City carried 476 containers? Let me recall: Gateway City (1957) - first purpose-built container ship, capacity 476 TEU? Some sources say 226. I'll avoid specific numbers I'm unsure of.)\n- Modern megaships: MSC G\u00fcls\u00fcn class ~23,000+ TEU; Ever Given (20,000 TEU)", "greedy": [{"prompt": "What is 17*23? Answer with just the number.", "tokens": [" No", " steps", ".\n\n", "What", " is", " ", "17", "*", "23", "?", " Answer", " with", " just", " the", " number", ".", " No", " steps", ".\n\n", "To", " solve", " the", " problem", ",", " I", " need", " to", " multiply", " ", "17", " by", " ", "23", ".\n\n", "17", " \u00d7", " ", "23", " =", " ", "39", "1", "\n\n", "Let", " me", " verify", ":", " ", "17", " \u00d7"], "token_logprobs": [-0.5339645147323608, -1.4312629699707031, -0.6227219700813293, -2.394000291824341, -0.034733857959508896, -0.07141785323619843, -0.12940452992916107, -0.04408905282616615, -0.0023861052468419075, -0.020640112459659576, -0.013627002947032452, -0.0008636558777652681, -0.000248043768806383, -0.00023779425828251988, -0.00016723664884921163, -0.005609011277556419, -0.003874415997415781, -0.00014375607133843005, -0.04082115367054939, -1.8987113237380981, -0.7420971393585205, -0.4911886751651764, -0.329582154750824, -0.4370775818824768, -0.7819880247116089, -0.30054205656051636, -0.0019647832959890366, -0.6340093612670898, -0.011875526048243046, -0.0002965487365145236, -0.03464646264910698, -1.8358061424805783e-05, -9.417489309271332e-06, -0.7518254518508911, -0.728101372718811, -0.6244921088218689, -0.00025507534155622125, -0.046296559274196625, -0.23665450513362885, -0.39430683851242065, -0.2540804445743561, -0.00014840454969089478, -0.16212260723114014, -1.3863519430160522, -0.020987454801797867, -0.3507349193096161, -0.8330035209655762, -0.12897050380706787, -0.11035574972629547, -0.004585228394716978]}, {"prompt": "Write a Python function that returns the nth Fibonacci number iteratively.", "tokens": [" The", " function", " should", " handle", " edge", " cases", " like", " n", "=", "0", " and", " n", "=", "1", ",", " and", " raise", " a", " ValueError", " for", " negative", " inputs", ".\n\n", "-", " Input", ":", " An", " integer", " n", "\n", "-", " Output", ":", " The", " nth", " Fibonacci", " number", "\n", "-", " Example", ":", " fib", "(", "10", ")", " should", " return", " ", "55", "\n"], "token_logprobs": [-1.9686455726623535, -0.502171516418457, -0.08533689379692078, -0.7450280785560608, -1.3920135498046875, -0.0012386515736579895, -1.299938678741455, -0.2477867603302002, -0.756680965423584, -0.004400332923978567, -0.21051783859729767, -0.06387461721897125, -0.0011737607419490814, -0.0001958416833076626, -1.0472832918167114, -0.2903238832950592, -0.6546998023986816, -0.15123315155506134, -0.008453183807432652, -0.314119428396225, -0.00752389058470726, -0.25283968448638916, -0.6313302516937256, -1.9620625972747803, -1.3374577760696411, -0.0016930069541558623, -0.1729714572429657, -0.0005467826849780977, -0.027555905282497406, -0.9808361530303955, -3.242440288886428e-05, -0.001444369088858366, -7.676783570786938e-05, -0.006336478982120752, -0.0012562725460156798, -9.786603914108127e-05, -3.099393507000059e-05, -0.5434079766273499, -0.002793697640299797, -0.3880147933959961, -0.029093634337186813, -0.6792575716972351, -0.0038443722296506166, -0.8405836820602417, -3.8742269680369645e-05, -0.1057451143860817, -4.815939246327616e-05, -2.0503786799963564e-05, -0.00016258825780823827, -0.4993346333503723]}, {"prompt": "Explain the difference between TCP and UDP in two sentences.", "tokens": [" Then", ",", " explain", " the", " difference", " between", " TCP", " and", " UDP", " in", " two", " sentences", ".", " Then", ",", " explain", " the", " difference", " between", " TCP", " and", " UDP", " in", " two", " sentences", ".", " Then", ",", " explain", " the", " difference", " between", " TCP", " and", " UDP", " in", " two", " sentences", ".", " Then", ",", " explain", " the", " difference", " between", " TCP", " and", " UDP", " in", " two"], "token_logprobs": [-1.722210168838501, -0.7200728058815002, -1.89879310131073, -1.7212132215499878, -0.3436768352985382, -0.16719898581504822, -1.5290614366531372, -0.027885088697075844, -0.07919464260339737, -0.21712860465049744, -0.650004506111145, -0.9500007033348083, -0.8788373470306396, -0.8283299207687378, -0.027243204414844513, -0.14736884832382202, -0.15356376767158508, -0.030464008450508118, -0.043157026171684265, -0.07051052153110504, -0.002471370855346322, -0.006156172137707472, -0.025828909128904343, -0.0355299636721611, -0.007403082679957151, -0.4034586250782013, -0.19530938565731049, -0.021165691316127777, -0.03506310284137726, -0.026420271024107933, -0.0041279117576777935, -0.009358584880828857, -0.00559620838612318, -0.001938013592734933, -0.0023844402749091387, -0.007875699549913406, -0.007675441447645426, -0.0027322364039719105, -0.2937863767147064, -0.07837116718292236, -0.006412169896066189, -0.00998761784285307, -0.017660275101661682, -0.004194747656583786, -0.004386090207844973, -0.004664492793381214, -0.0019645453430712223, -0.0017496529035270214, -0.005039251875132322, -0.002457695547491312]}, {"prompt": "Translate to French: 'The quick brown fox jumps over the lazy dog.'", "tokens": [" The", " quick", " brown", " fox", " jumps", " over", " the", " lazy", " dog", ".\n\n", "The", " quick", " brown", " fox", " jumps", " over", " the", " lazy", " dog", ".\n\n", "The", " quick", " brown", " fox", " jumps", " over", " the", " lazy", " dog", ".\n\n", "The", " quick", " brown", " fox", " jumps", " over", " the", " lazy", " dog", ".\n\n", "The", " quick", " brown", " fox", " jumps", " over", " the", " lazy", " dog", ".\n\n"], "token_logprobs": [-1.649539589881897, -0.08263400197029114, -0.0029683122411370277, -0.0024340308737009764, -0.021369799971580505, -0.0005610798834823072, -0.0008338788175024092, -0.0007481159991584718, -0.0010244365548714995, -0.7542631030082703, -0.7775561809539795, -0.6337171196937561, -0.002058055717498064, -0.0024215441662818193, -0.0024176198057830334, -9.7508447652217e-05, -0.0004067785630468279, -0.0029261175077408552, -0.00021395778458099812, -0.47863155603408813, -0.2845117151737213, -0.08507601916790009, -0.005965760443359613, -0.002583144698292017, -0.0025943215005099773, -6.603976362384856e-05, -0.0004301338049117476, -0.0006644901586696506, -0.001312705222517252, -0.16103971004486084, -0.22266927361488342, -0.028461331501603127, -0.0022077015601098537, -0.0040604774840176105, -0.000764792668633163, -0.00025948495022021234, -0.00018594920402392745, -0.00014673586701974273, -0.0006422125734388828, -0.07882342487573624, -0.26711925864219666, -0.037533245980739594, -0.001487697591073811, -0.004707681946456432, -0.001576372655108571, -0.0003116837178822607, -0.0005117537220939994, -0.0021822466515004635, -0.0002740246127359569, -0.15439195930957794]}, {"prompt": "List the first 8 prime numbers separated by commas.", "tokens": [" ", "2", ",", "3", ",", "5", ",", "7", ",", "11", ",", "13", ",", "17", ",", "19", "\n\n", "F", "ilo", " tutor", " solution", "\n\n", "Students", " who", " ask", " this", " question", " also", " asked", "\n\n", "|", " Question", " Text", " ", " |", " The", " first", " ", "8", " prime", " numbers", " separated", " by", " commas", " are", ":", " ", "2", ",", " "], "token_logprobs": [-2.47273325920105, -0.8465018272399902, -0.20934341847896576, -0.7100750803947449, -0.009551658295094967, -0.010330774821341038, -0.03215751424431801, -0.021142814308404922, -0.04794848710298538, -0.054938048124313354, -0.018736518919467926, -0.009158269502222538, -0.06906037032604218, -0.03428402170538902, -0.10199463367462158, -0.032381437718868256, -1.1714547872543335, -1.9063001871109009, -0.0005490464391186833, -0.035503387451171875, -0.006995352916419506, -4.5298504119273275e-05, -0.007928685285151005, -0.0001629458274692297, -9.440929716220126e-05, -3.576278118089249e-07, -1.1920928244535389e-07, -1.0728830375228426e-06, -0.003818127792328596, -0.0006684214458800852, -0.041534896939992905, -8.940656698541716e-06, -6.69933797325939e-05, -1.4305104514278355e-06, -0.000286657246761024, -0.20002372562885284, -0.004902722779661417, -0.0014604389434680343, -0.0007986216805875301, -0.0005770448478870094, -4.9470632802695036e-05, -0.5132142305374146, -9.83428253675811e-05, -0.0007053509471006691, -0.39528897404670715, -0.7401394844055176, -0.11407836526632309, -0.0004741021548397839, -0.0009533389820717275, -0.1799558699131012]}, {"prompt": "What year did the Apollo 11 mission land on the moon? One word answer.", "tokens": [" ", "196", "9", ".", " What", " was", " the", " name", " of", " the", " first", " artificial", " Earth", " satellite", "?", " S", "put", "nik", ".", " What", " is", " the", " name", " of", " the", " largest", " planet", " in", " our", " solar", " system", "?", " Jupiter", ".", " What", " is", " the", " name", " of", " the", " NASA", " space", " telescope", " launched", " in", " ", "202", "1", " that", " is"], "token_logprobs": [-2.9569573402404785, -0.5128473043441772, -0.0327749103307724, -0.9968858957290649, -1.7308241128921509, -1.5436567068099976, -0.06314665824174881, -0.5276843309402466, -0.01302390731871128, -0.04848307743668556, -0.6253899335861206, -1.4379539489746094, -0.39747557044029236, -0.008306360803544521, -0.3908175230026245, -0.8787673115730286, -0.0015930355293676257, -0.0002949994814116508, -0.14244797825813293, -1.1112890243530273, -0.6121114492416382, -0.03419646993279457, -0.8894662857055664, -0.006880045868456364, -0.5993072986602783, -0.985981822013855, -0.0672343522310257, -0.00864111352711916, -0.009636197239160538, -0.02014872245490551, -0.001007764250971377, -0.02126244083046913, -0.03388837352395058, -0.09148004651069641, -0.46715137362480164, -0.3278219401836395, -0.045866355299949646, -0.39599132537841797, -0.005201378837227821, -0.5893699526786804, -0.8430454730987549, -0.49212121963500977, -0.03188123553991318, -0.2139008790254593, -0.06327691674232483, -0.023227160796523094, -0.5247340202331543, -0.0011635207338258624, -0.5196314454078674, -0.9678736925125122]}, {"prompt": "Summarize the plot of Romeo and Juliet in one sentence.", "tokens": [" Then", ",", " summarize", " the", " plot", " of", " another", " familiar", " story", " in", " one", " sentence", ".", " Compare", " the", " two", " sentences", ".", " What", " do", " they", " have", " in", " common", "?", " What", " is", " different", "?", " What", " do", " these", " similarities", " and", " differences", " suggest", " about", " the", " nature", " of", " narrative", "?\n\n", "The", " chronological", " order", " of", " events", " related", " by", " the"], "token_logprobs": [-2.510852098464966, -0.38090285658836365, -1.6050550937652588, -0.34692907333374023, -0.2790347635746002, -0.27269163727760315, -1.7477823495864868, -1.3371325731277466, -0.022476447746157646, -0.43683677911758423, -0.4575590193271637, -0.026687266305088997, -0.42049771547317505, -2.146791458129883, -0.7664984464645386, -0.5831676125526428, -0.9063276052474976, -1.1274220943450928, -0.742306113243103, -1.5807318687438965, -0.463785320520401, -0.1954614520072937, -0.0026043090038001537, -0.03041706047952175, -0.23723043501377106, -0.8948622345924377, -1.7165815830230713, -0.287236750125885, -0.6562533378601074, -1.3421944379806519, -1.4933453798294067, -0.596420168876648, -0.3422274887561798, -0.04228927195072174, -0.005274428054690361, -0.5505183935165405, -0.3697821795940399, -0.4134998619556427, -1.8081897497177124, -0.09521330893039703, -1.6461513042449951, -1.093263864517212, -2.232003688812256, -1.009954810142517, -2.610649426060263e-05, 0.0, -1.0847986231965479e-05, -0.05477927252650261, -0.038065169006586075, -1.311301275563892e-06]}, {"prompt": "Write a SQL query selecting the top 5 customers by total order value from tables customers(id,name) and orders(id,customer_id,value).", "tokens": [" Show", " customer", " name", " and", " total", " value", ".\n\n", "SQL", " Query", ":", " Top", " ", "5", " Customers", " by", " Total", " Order", " Value", "\n\n", "This", " SQL", " query", " retrieves", " the", " top", " ", "5", " customers", " based", " on", " their", " total", " order", " value", ",", " joining", " the", " `", "customers", " and", "`", " and", " `", "orders", " tables", " and", " aggreg", "ating", " order", " values"], "token_logprobs": [-1.2284963130950928, -0.8918663263320923, -0.09108343720436096, -0.05327378958463669, -0.1501082479953766, -0.8929679989814758, -1.0398002862930298, -1.1595065593719482, -1.0504779815673828, -0.8854727745056152, -0.11332081258296967, -0.018643619492650032, -5.5549986427649856e-05, -0.001874001114629209, -0.008812228217720985, -0.013624768704175949, -0.0061508407816290855, -0.0003083468764089048, -0.05783206969499588, -1.3851438760757446, -1.4397640228271484, -0.04293013736605644, -0.5352510213851929, -0.000497336674015969, -0.01697123982012272, -0.0012138864258304238, -1.2278481335670222e-05, -0.0006150499684736133, -0.4459102153778076, -2.4676019165781327e-05, -0.018381137400865555, -0.005231619346886873, -0.0026886051055043936, -0.0031163957901299, -1.1535018682479858, -0.8411168456077576, -0.3593292832374573, -0.21669110655784607, -0.0029645089525729418, -0.3164861798286438, -0.0003890234511345625, -0.0002964295563288033, -9.655486064730212e-05, -0.000940714729949832, -0.004669950809329748, -1.24745774269104, -0.4944056570529938, -1.5497195136049413e-06, -0.5822879672050476, -0.17569199204444885]}, {"prompt": "What is the derivative of x^3 + 2x with respect to x?", "tokens": [" A", ")", " ", "3", "x", "^", "2", " +", " ", "2", " B", ")", " ", "3", "x", "^", "2", " C", ")", " x", "^", "2", " +", " ", "2", " D", ")", " ", "3", "x", " +", " ", "2", "\n\n", "Der", "ivative", " of", " x", "^", "3", " +", " ", "2", "x", " is", " ", "3", "x", "^", "2"], "token_logprobs": [-1.1993834972381592, -0.08901525288820267, -0.04353532940149307, -0.00991373136639595, -0.014755168929696083, -0.026040781289339066, -0.005995621904730797, -0.29368090629577637, -0.09320046752691269, -0.0017926113214343786, -0.6470931768417358, -0.014025573618710041, -0.17031075060367584, -0.04926355183124542, -0.004726903047412634, -0.1721748560667038, -0.03880072012543678, -0.8462035059928894, -0.0010372978867962956, -0.47493329644203186, -0.0017999890260398388, -0.024153385311365128, -0.011953625828027725, -0.0003997480380348861, -0.0014006814453750849, -0.03701401874423027, -0.00167598866391927, -0.07545437663793564, -0.5835696458816528, -0.002371596172451973, -0.7554231882095337, -0.0011936451774090528, -0.0008511496125720441, -0.047347553074359894, -1.9332166910171509, -0.01650102809071541, -0.08500517904758453, -0.2852739691734314, -0.05047726631164551, -0.020355843007564545, -0.32135993242263794, -0.0003228858404327184, -4.482168878894299e-05, -0.00020203932945150882, -0.5257324576377869, -0.39029359817504883, -0.0006705658161081374, -0.0002648479712661356, -0.016299204900860786, -3.7431014789035544e-05]}, {"prompt": "Name the chemical symbol for gold and the element with atomic number 6.", "tokens": [" Also", ",", " what", " is", " the", " chemical", " symbol", " for", " sodium", "?\n\n", "Under", "stand", " the", " Problem", "\n\n", "The", " question", " is", " asking", " for", " the", " chemical", " symbol", " for", " gold", ",", " the", " element", " with", " atomic", " number", " ", "6", ",", " and", " the", " chemical", " symbol", " for", " sodium", ".", " This", " involves", " basic", " chemistry", " knowledge", ".\n\n", "Answer", "\n\n", "Gold"], "token_logprobs": [-1.858495831489563, -0.16975700855255127, -1.756112813949585, -0.42193692922592163, -0.20150718092918396, -1.3564766645431519, -0.3985130488872528, -0.061088450253009796, -1.7535368204116821, -0.6411986351013184, -0.6826284527778625, -0.004277010448276997, -0.0001370812824461609, -0.0012516292044892907, -7.629365427419543e-06, -0.001022888463921845, -0.02649538405239582, -0.4306084215641022, -0.009948669001460075, -0.1586037129163742, -0.30859947204589844, -0.006336716003715992, -0.43078285455703735, -0.5260207653045654, -0.041210636496543884, -0.1543508917093277, -0.09545984864234924, -0.4936869740486145, -0.028597230091691017, -0.008440890349447727, -3.802703940891661e-05, -7.903263758635148e-05, -0.00012742661056108773, -0.017153475433588028, -0.041007332503795624, -0.015617916360497475, -0.025308450683951378, -0.0005539313424378633, -4.8636207793606445e-05, -0.0005229535745456815, -0.2904960811138153, -0.5248113870620728, -0.9419708847999573, -0.7357808351516724, -0.18394440412521362, -0.0033896868117153645, -1.0488858222961426, -9.440929716220126e-05, -3.0874729418428615e-05, -0.301129549741745]}], "scored": [{"text": "The mitochondrion is the powerhouse of the cell, converting ", "prompt_logprobs": [null, {"53582": {"logprob": -11.927324295043945, "rank": 20447}, "154822": {"logprob": -4.3648247718811035, "rank": 1}}, {"81": {"logprob": -1.1411426067352295, "rank": 2}, "4204": {"logprob": -0.3911426365375519, "rank": 1}}, {"290": {"logprob": -7.510157047363464e-06, "rank": 1}}, {"374": {"logprob": -0.4827197790145874, "rank": 1}}, {"279": {"logprob": -1.0389354228973389, "rank": 1}}, {"73538": {"logprob": -3.3948867321014404, "rank": 3}, "1240": {"logprob": -0.26988667249679565, "rank": 1}}, {"315": {"logprob": -0.0378556028008461, "rank": 1}}, {"279": {"logprob": -0.043669652193784714, "rank": 1}}, {"2779": {"logprob": -0.04398604482412338, "rank": 1}}, {"11": {"logprob": -1.3363630771636963, "rank": 2}, "13": {"logprob": -0.9613630175590515, "rank": 1}}, {"33277": {"logprob": -3.7138848304748535, "rank": 10}, "323": {"logprob": -1.4013848304748535, "rank": 1}}, {"36209": {"logprob": -1.0502684116363525, "rank": 1}}, {"1119": {"logprob": -0.07778626680374146, "rank": 1}}, {"993": {"logprob": -3.8446688652038574, "rank": 4}, "66072": {"logprob": -0.46966883540153503, "rank": 1}}, {"70338": {"logprob": -6.48477507638745e-05, "rank": 1}}, {"482": {"logprob": -7.080780778778717e-05, "rank": 1}}, {"2406": {"logprob": -0.0011192255187779665, "rank": 1}}, {"759": {"logprob": -0.006047166883945465, "rank": 1}}, {"91597": {"logprob": -0.0001668790791882202, "rank": 1}}, {"1526": {"logprob": -4.401804447174072, "rank": 2}, "320": {"logprob": -0.026804490014910698, "rank": 1}}, {"77673": {"logprob": -0.8898406028747559, "rank": 1}}, {"93199": {"logprob": -0.004327219445258379, "rank": 1}}, {"2302": {"logprob": -0.00014220656885299832, "rank": 1}}, {"13": {"logprob": -0.6309416890144348, "rank": 1}}, {"1096": {"logprob": -2.43224835395813, "rank": 3}, "21714": {"logprob": -1.2447483539581299, "rank": 1}}, {"1882": {"logprob": -0.6010481715202332, "rank": 1}}, {"13657": {"logprob": -4.120265483856201, "rank": 9}, "33482": {"logprob": -1.2452654838562012, "rank": 1}}, {"3941": {"logprob": -0.9921097159385681, "rank": 1}}, {"279": {"logprob": -0.3731870651245117, "rank": 1}}, {"9176": {"logprob": -0.1444033980369568, "rank": 1}}, {"70428": {"logprob": -0.2330683320760727, "rank": 1}}, {"38346": {"logprob": -0.06355483829975128, "rank": 1}}, {"11": {"logprob": -0.9225961565971375, "rank": 1}}, {"1380": {"logprob": -1.8594939708709717, "rank": 2}, "892": {"logprob": -0.8594939708709717, "rank": 1}}, {"279": {"logprob": -0.5926007032394409, "rank": 1}}, {"16698": {"logprob": -0.30325180292129517, "rank": 1}}, {"7557": {"logprob": -0.006079395767301321, "rank": 1}}, {"8780": {"logprob": -0.0015979153104126453, "rank": 1}}, {"63111": {"logprob": -1.9291372299194336, "rank": 2}, "323": {"logprob": -1.6791372299194336, "rank": 1}}, {"264": {"logprob": -0.09616586565971375, "rank": 1}}, {"80822": {"logprob": -0.005951184779405594, "rank": 1}}, {"20129": {"logprob": -0.13866667449474335, "rank": 1}}, {"13": {"logprob": -1.9624547958374023, "rank": 2}, "429": {"logprob": -0.5874547362327576, "rank": 1}}]}, {"text": "def quicksort(arr):\n if len(arr) <= 1:\n return arr", "prompt_logprobs": [null, {"3974": {"logprob": -7.5163984298706055, "rank": 187}, "1159": {"logprob": -3.4538986682891846, "rank": 1}}, {"6860": {"logprob": -0.26104360818862915, "rank": 1}}, {"10934": {"logprob": -0.4579106569290161, "rank": 1}}, {"982": {"logprob": -0.7912289500236511, "rank": 2}, "11": {"logprob": -0.6662289500236511, "rank": 1}}, {"262": {"logprob": -0.09921220690011978, "rank": 1}}, {"421": {"logprob": -0.2923320531845093, "rank": 1}}, {"2422": {"logprob": -0.03586964309215546, "rank": 1}}, {"10934": {"logprob": -0.0004549183649942279, "rank": 1}}, {"8": {"logprob": -0.02526521123945713, "rank": 1}}, {"2651": {"logprob": -0.1488965004682541, "rank": 1}}, {"220": {"logprob": -0.009869704023003578, "rank": 1}}, {"16": {"logprob": -0.0006195771275088191, "rank": 1}}, {"510": {"logprob": -0.010196853429079056, "rank": 1}}, {"286": {"logprob": -0.0006637753685936332, "rank": 1}}, {"470": {"logprob": -0.00012194366718176752, "rank": 1}}, {"2890": {"logprob": -0.00046564225340262055, "rank": 1}}, {"198": {"logprob": -0.05762660503387451, "rank": 1}}, {"262": {"logprob": -0.001563875237479806, "rank": 1}}, {"25964": {"logprob": -0.3331810534000397, "rank": 1}}, {"284": {"logprob": -0.00428733741864562, "rank": 1}}, {"2890": {"logprob": -0.00788823701441288, "rank": 1}}, {"24617": {"logprob": -0.2954382300376892, "rank": 1}}, {"10934": {"logprob": -6.806619057897478e-05, "rank": 1}}, {"8": {"logprob": -0.10586930066347122, "rank": 1}}, {"442": {"logprob": -0.04571877792477608, "rank": 1}}, {"220": {"logprob": -0.00044860312482342124, "rank": 1}}, {"17": {"logprob": -0.0008239926537498832, "rank": 1}}, {"921": {"logprob": -0.022608501836657524, "rank": 1}}, {"262": {"logprob": -0.0009258274803869426, "rank": 1}}, {"2115": {"logprob": -0.07724074274301529, "rank": 1}}, {"284": {"logprob": -0.0031280419789254665, "rank": 1}}, {"508": {"logprob": -0.01961255632340908, "rank": 1}}, {"87": {"logprob": -0.0030618475284427404, "rank": 1}}, {"369": {"logprob": -0.0007338214782066643, "rank": 1}}, {"856": {"logprob": -1.7165990357170813e-05, "rank": 1}}, {"304": {"logprob": -4.172316494077677e-06, "rank": 1}}, {"2890": {"logprob": -4.672895011026412e-05, "rank": 1}}, {"421": {"logprob": -0.0003597089380491525, "rank": 1}}, {"856": {"logprob": -0.00037353215157054365, "rank": 1}}, {"366": {"logprob": -0.0003488647344056517, "rank": 1}}, {"25964": {"logprob": -0.00014852374442853034, "rank": 1}}, {"921": {"logprob": -0.014500945806503296, "rank": 1}}, {"262": {"logprob": -4.756337511935271e-05, "rank": 1}}, {"6149": {"logprob": -0.04473856836557388, "rank": 1}}, {"284": {"logprob": -0.0003844952443614602, "rank": 1}}, {"508": {"logprob": -0.00013457823661156, "rank": 1}}, {"87": {"logprob": -6.294052582234144e-05, "rank": 1}}, {"369": {"logprob": -6.758938252460212e-05, "rank": 1}}, {"856": {"logprob": -4.768360213347478e-06, "rank": 1}}, {"304": {"logprob": -6.198863957251888e-06, "rank": 1}}, {"2890": {"logprob": -7.986990567587782e-06, "rank": 1}}, {"421": {"logprob": -0.00016282663273159415, "rank": 1}}, {"856": {"logprob": -2.2172682292875834e-05, "rank": 1}}, {"621": {"logprob": -0.00026639728457666934, "rank": 1}}, {"25964": {"logprob": -0.00011646069469861686, "rank": 1}}, {"921": {"logprob": -0.002465306082740426, "rank": 1}}, {"262": {"logprob": -5.6622808187967166e-05, "rank": 1}}, {"1290": {"logprob": -0.0009413101943209767, "rank": 1}}, {"284": {"logprob": -0.0001113352773245424, "rank": 1}}, {"508": {"logprob": -0.0007493072189390659, "rank": 1}}, {"87": {"logprob": -5.9960475482512265e-05, "rank": 1}}, {"369": {"logprob": -3.957670196541585e-05, "rank": 1}}, {"856": {"logprob": -8.34461570775602e-06, "rank": 1}}, {"304": {"logprob": -4.768370445162873e-07, "rank": 1}}, {"2890": {"logprob": -3.0636318115284666e-05, "rank": 1}}, {"421": {"logprob": -0.0001656871900195256, "rank": 1}}, {"856": {"logprob": -2.288792165927589e-05, "rank": 1}}, {"861": {"logprob": -0.00021479207498487085, "rank": 1}}, {"25964": {"logprob": -3.7788631743751466e-05, "rank": 1}}, {"921": {"logprob": -0.012385360896587372, "rank": 1}}, {"262": {"logprob": -0.004069263115525246, "rank": 1}}, {"470": {"logprob": -0.0049212281592190266, "rank": 1}}, {"3974": {"logprob": -0.001862459466792643, "rank": 1}}, {"6860": {"logprob": -4.31528314948082e-05, "rank": 1}}, {"17646": {"logprob": -0.0008093419019132853, "rank": 1}}, {"8": {"logprob": -0.00011789103882620111, "rank": 1}}, {"488": {"logprob": -2.4318398573086597e-05, "rank": 1}}, {"6149": {"logprob": -0.008893049322068691, "rank": 1}}, {"488": {"logprob": -5.030505417380482e-05, "rank": 1}}, {"3974": {"logprob": -0.0016149348812177777, "rank": 1}}, {"6860": {"logprob": -7.807903602952138e-05, "rank": 1}}, {"27611": {"logprob": -0.0006333967321552336, "rank": 1}}, {"8": {"logprob": -3.2184078693389893, "rank": 4}, "692": {"logprob": -0.3434078097343445, "rank": 1}}]}, {"text": "In 1969, the Apollo 11 mission successfully landed the first", "prompt_logprobs": [null, {"220": {"logprob": -6.883457660675049, "rank": 32}, "154822": {"logprob": -2.633457660675049, "rank": 1}}, {"121818": {"logprob": -6.685415744781494, "rank": 75}, "17": {"logprob": -1.6229157447814941, "rank": 1}}, {"24": {"logprob": -2.0959348678588867, "rank": 2}, "17": {"logprob": -1.9709348678588867, "rank": 1}}, {"11": {"logprob": -0.3142923414707184, "rank": 1}}, {"279": {"logprob": -1.8265211582183838, "rank": 1}}, {"34976": {"logprob": -2.9791653156280518, "rank": 2}, "356": {"logprob": -1.6666653156280518, "rank": 1}}, {"220": {"logprob": -0.29257044196128845, "rank": 1}}, {"98965": {"logprob": -0.1530940681695938, "rank": 1}}, {"8951": {"logprob": -1.0149643421173096, "rank": 1}}, {"7790": {"logprob": -2.837709426879883, "rank": 5}, "26039": {"logprob": -1.1502095460891724, "rank": 1}}, {"26039": {"logprob": -0.18243205547332764, "rank": 1}}, {"279": {"logprob": -1.460514783859253, "rank": 2}, "12671": {"logprob": -0.9605148434638977, "rank": 1}}, {"1156": {"logprob": -0.013778219930827618, "rank": 1}}, {"12671": {"logprob": -0.1814669966697693, "rank": 1}}, {"389": {"logprob": -0.012509924359619617, "rank": 1}}, {"279": {"logprob": -0.0027726562693715096, "rank": 1}}, {"17309": {"logprob": -0.5299549102783203, "rank": 1}}, {"13": {"logprob": -0.41980794072151184, "rank": 1}}, {"32962": {"logprob": -1.7579048871994019, "rank": 2}, "1096": {"logprob": -1.2579048871994019, "rank": 1}}, {"44605": {"logprob": -0.0029639145359396935, "rank": 1}}, {"323": {"logprob": -0.2153635174036026, "rank": 1}}, {"37754": {"logprob": -0.06758884340524673, "rank": 1}}, {"30230": {"logprob": -0.00026055757189169526, "rank": 1}}, {"25210": {"logprob": -8.654219709569588e-05, "rank": 1}}, {"7391": {"logprob": -2.743393659591675, "rank": 4}, "6116": {"logprob": -0.7433936595916748, "rank": 1}}, {"13179": {"logprob": -2.4915246963500977, "rank": 4}, "220": {"logprob": -1.2415248155593872, "rank": 1}}, {"1378": {"logprob": -3.5392518043518066, "rank": 2}, "220": {"logprob": -0.03925185278058052, "rank": 1}}, {"323": {"logprob": -0.19062088429927826, "rank": 1}}, {"264": {"logprob": -0.010297974571585655, "rank": 1}}, {"8337": {"logprob": -8.500765800476074, "rank": 3}, "4279": {"logprob": -0.000766102981287986, "rank": 1}}, {"4115": {"logprob": -0.004973184317350388, "rank": 1}}, {"4889": {"logprob": -2.118410348892212, "rank": 3}, "389": {"logprob": -0.6184103488922119, "rank": 1}}, {"279": {"logprob": -0.1295350342988968, "rank": 1}}, {"41305": {"logprob": -0.27623507380485535, "rank": 1}}, {"11": {"logprob": -0.8422405123710632, "rank": 2}, "1393": {"logprob": -0.7172405123710632, "rank": 1}}, {"25814": {"logprob": -2.261660575866699, "rank": 3}, "1393": {"logprob": -0.5116605162620544, "rank": 1}}, {"56329": {"logprob": -0.9502583742141724, "rank": 1}}, {"3684": {"logprob": -2.332808494567871, "rank": 2}, "10464": {"logprob": -0.20780859887599945, "rank": 1}}, {"311": {"logprob": -3.766075372695923, "rank": 4}, "323": {"logprob": -0.14107537269592285, "rank": 1}}, {"4446": {"logprob": -0.010382096283137798, "rank": 1}}, {"1182": {"logprob": -0.0022718114778399467, "rank": 1}}, {"311": {"logprob": -0.028401080518960953, "rank": 1}}, {"9234": {"logprob": -0.001636000582948327, "rank": 1}}, {"13": {"logprob": -0.3184300661087036, "rank": 1}}]}, {"text": "Le petit prince demanda au renard ce que signifiait le mot a", "prompt_logprobs": [null, {"44744": {"logprob": -22.625015258789062, "rank": 21419}, "154822": {"logprob": -1.537788011773955e-05, "rank": 1}}, {"41490": {"logprob": -2.5846405029296875, "rank": 1}}, {"137474": {"logprob": -12.593887329101562, "rank": 6032}, "284": {"logprob": -2.4376368522644043, "rank": 1}}, {"7906": {"logprob": -4.0444746017456055, "rank": 7}, "34327": {"logprob": -0.48197484016418457, "rank": 1}}, {"5672": {"logprob": -5.072017192840576, "rank": 17}, "11454": {"logprob": -0.9470173120498657, "rank": 1}}, {"567": {"logprob": -0.0007189311436377466, "rank": 1}}, {"3761": {"logprob": -5.452221870422363, "rank": 28}, "549": {"logprob": -1.8272217512130737, "rank": 1}}, {"1709": {"logprob": -0.3002340495586395, "rank": 1}}, {"1841": {"logprob": -1.2496821880340576, "rank": 2}, "272": {"logprob": -0.9996821880340576, "rank": 1}}, {"333": {"logprob": -0.018434623256325722, "rank": 1}}, {"685": {"logprob": -0.0025433117989450693, "rank": 1}}, {"275": {"logprob": -0.01985321193933487, "rank": 1}}, {"512": {"logprob": -1.861337661743164, "rank": 3}, "12480": {"logprob": -1.361337661743164, "rank": 1}}, {"3852": {"logprob": -0.11574337631464005, "rank": 1}}, {"131231": {"logprob": -2.995323657989502, "rank": 3}, "12480": {"logprob": -0.8703237175941467, "rank": 1}}, {"6496": {"logprob": -3.421248038648628e-05, "rank": 1}}, {"12053": {"logprob": -0.3106018900871277, "rank": 1}}, {"13": {"logprob": -1.166166067123413, "rank": 1}}, {"1967": {"logprob": -1.0800580978393555, "rank": 1}}, {"5672": {"logprob": -0.03638773784041405, "rank": 1}}, {"567": {"logprob": -0.00023767507809679955, "rank": 1}}, {"3247": {"logprob": -4.237571716308594, "rank": 4}, "24324": {"logprob": -0.2375718206167221, "rank": 1}}, {"5011": {"logprob": -0.0007352509419433773, "rank": 1}}, {"64": {"logprob": -0.0007251255447044969, "rank": 1}}, {"1709": {"logprob": -0.5658798813819885, "rank": 1}}, {"44244": {"logprob": -1.0300593376159668, "rank": 2}, "272": {"logprob": -0.7800593376159668, "rank": 1}}, {"1841": {"logprob": -0.6900672316551208, "rank": 1}}, {"333": {"logprob": -0.00031740395934320986, "rank": 1}}, {"685": {"logprob": -0.0029820995405316353, "rank": 1}}, {"275": {"logprob": -0.0007614573696628213, "rank": 1}}, {"1884": {"logprob": -9.713112831115723, "rank": 68}, "74145": {"logprob": -0.21311283111572266, "rank": 1}}, {"261": {"logprob": -0.0070532383397221565, "rank": 1}}, {"939": {"logprob": -0.0019070786656811833, "rank": 1}}, {"151101": {"logprob": -0.0020761380437761545, "rank": 1}}, {"11": {"logprob": -1.6812024116516113, "rank": 2}, "13": {"logprob": -0.6812024712562561, "rank": 1}}, {"1842": {"logprob": -1.5158237218856812, "rank": 1}}, {"1709": {"logprob": -0.980962872505188, "rank": 1}}, {"4403": {"logprob": -3.8899288177490234, "rank": 10}, "1884": {"logprob": -1.5149288177490234, "rank": 1}}, {"512": {"logprob": -1.3587359189987183, "rank": 1}}, {"41490": {"logprob": -2.84769868850708, "rank": 2}, "44744": {"logprob": -0.09769879281520844, "rank": 1}}, {"326": {"logprob": -0.6537078022956848, "rank": 1}}, {"6": {"logprob": -0.15203288197517395, "rank": 1}}, {"138865": {"logprob": -0.036557964980602264, "rank": 1}}, {"6496": {"logprob": -0.011730148456990719, "rank": 1}}, {"285": {"logprob": -0.00015269544383045286, "rank": 1}}, {"1315": {"logprob": -0.000251142424531281, "rank": 1}}, {"11": {"logprob": -0.04195988178253174, "rank": 1}}, {"44786": {"logprob": -0.5310053825378418, "rank": 1}}, {"38832": {"logprob": -2.0455362796783447, "rank": 3}, "34467": {"logprob": -0.6705363392829895, "rank": 1}}, {"1167": {"logprob": -0.00026544384309090674, "rank": 1}}, {"62129": {"logprob": -0.07335690408945084, "rank": 1}}, {"326": {"logprob": -0.04246489331126213, "rank": 1}}, {"21997": {"logprob": -0.023533552885055542, "rank": 1}}, {"409": {"logprob": -0.08645187318325043, "rank": 1}}, {"326": {"logprob": -0.006858379580080509, "rank": 1}}, {"48052": {"logprob": -0.0011840007500723004, "rank": 1}}, {"265": {"logprob": -3.540453326422721e-05, "rank": 1}}, {"13": {"logprob": -2.4708456993103027, "rank": 2}, "1842": {"logprob": -0.22084560990333557, "rank": 1}}]}, {"text": "The gradient of the loss function with respect to the weight", "prompt_logprobs": [null, {"20129": {"logprob": -8.958574295043945, "rank": 992}, "154822": {"logprob": -4.3648247718811035, "rank": 1}}, {"315": {"logprob": -0.7530007362365723, "rank": 1}}, {"279": {"logprob": -1.0023329257965088, "rank": 2}, "264": {"logprob": -0.7523329854011536, "rank": 1}}, {"4709": {"logprob": -4.664595603942871, "rank": 11}, "729": {"logprob": -1.039595365524292, "rank": 1}}, {"729": {"logprob": -0.26539698243141174, "rank": 1}}, {"448": {"logprob": -2.0058484077453613, "rank": 2}, "374": {"logprob": -1.3808484077453613, "rank": 1}}, {"5091": {"logprob": -0.008570555597543716, "rank": 1}}, {"311": {"logprob": -0.000613143783994019, "rank": 1}}, {"279": {"logprob": -0.49543672800064087, "rank": 1}}, {"14314": {"logprob": -0.9178521037101746, "rank": 1}}, {"374": {"logprob": -1.0224812030792236, "rank": 1}}, {"24113": {"logprob": -2.9625673294067383, "rank": 6}, "264": {"logprob": -1.0875673294067383, "rank": 1}}, {"4566": {"logprob": -3.365877151489258, "rank": 6}, "1667": {"logprob": -0.9908771514892578, "rank": 1}}, {"1182": {"logprob": -0.3806135952472687, "rank": 1}}, {"2674": {"logprob": -0.013360965996980667, "rank": 1}}, {"27048": {"logprob": -0.011677246540784836, "rank": 1}}, {"11": {"logprob": -1.0888934135437012, "rank": 1}}, {"18915": {"logprob": -4.6767144203186035, "rank": 10}, "892": {"logprob": -0.6767145991325378, "rank": 1}}, {"279": {"logprob": -0.007259064819663763, "rank": 1}}, {"8780": {"logprob": -0.006695337127894163, "rank": 1}}, {"5912": {"logprob": -0.0010518262861296535, "rank": 1}}, {"6193": {"logprob": -1.847823143005371, "rank": 2}, "1526": {"logprob": -0.8478232026100159, "rank": 1}}, {"553": {"logprob": -0.008238380774855614, "rank": 1}}, {"6193": {"logprob": -0.0006792622152715921, "rank": 1}}, {"504": {"logprob": -1.6093263626098633, "rank": 2}, "13": {"logprob": -0.8593264222145081, "rank": 1}}, {"279": {"logprob": -0.02686193771660328, "rank": 1}}, {"2550": {"logprob": -0.014062248170375824, "rank": 1}}, {"1182": {"logprob": -0.34703919291496277, "rank": 1}}, {"311": {"logprob": -0.05637872591614723, "rank": 1}}, {"279": {"logprob": -0.0033738852944225073, "rank": 1}}, {"1946": {"logprob": -0.013425423763692379, "rank": 1}}, {"13": {"logprob": -0.522415041923523, "rank": 1}}, {"794": {"logprob": -7.078248977661133, "rank": 51}, "1752": {"logprob": -1.328249216079712, "rank": 1}}, {"65474": {"logprob": -0.01044945977628231, "rank": 1}}, {"20129": {"logprob": -1.9303635358810425, "rank": 2}, "52755": {"logprob": -0.18036355078220367, "rank": 1}}, {"36760": {"logprob": -0.005931513383984566, "rank": 1}}, {"1221": {"logprob": -3.7587532997131348, "rank": 5}, "320": {"logprob": -0.44625324010849, "rank": 1}}, {"8836": {"logprob": -0.07871434837579727, "rank": 1}}, {"1817": {"logprob": -1.2432700395584106, "rank": 2}, "279": {"logprob": -0.36827003955841064, "rank": 1}}, {"4680": {"logprob": -0.4352835714817047, "rank": 1}}, {"21070": {"logprob": -6.724987983703613, "rank": 32}, "553": {"logprob": -1.4124879837036133, "rank": 1}}, {"745": {"logprob": -0.0024585279170423746, "rank": 1}}, {"13": {"logprob": -7.883650779724121, "rank": 8}, "311": {"logprob": -0.008650803938508034, "rank": 1}}]}]}