GLM-5.3-Flash-NVFP4-FP8ATTN-512K / scripts /confirm-8666-512k.log
tacos4me's picture
evidence: TF@8666 dump, 8666@512k confirmation log, convert_fp8visual.py
6bd1d23 verified
Raw
History Blame Contribute Delete
68.7 kB
[budget] waiting for GPU lock...
[budget] holding GPU lock
(APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345]
(APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345] █ █ █▄ ▄█
(APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.1.dev20051+g487ecf187
(APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345] █▄█▀ █ █ █ █ model /model
(APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:345]
(APIServer pid=59) INFO 08-30 23:38:57 [api_utils.py:273] non-default args: {'model_tag': '/model', 'enable_auto_tool_choice': True, 'tool_call_parser': 'glm47', 'host': '0.0.0.0', 'port': 18998, 'model': '/model', 'trust_remote_code': True, 'max_model_len': 524288, 'served_model_name': ['glm-max'], 'reasoning_parser': 'glm45', 'tensor_parallel_size': 2, 'gpu_memory_utilization': 0.95, 'kv_cache_memory_bytes': 3484000000, 'kv_cache_dtype': 'fp8', 'enable_prefix_caching': True, 'limit_mm_per_prompt': {'image': 1, 'video': 0}, 'max_num_batched_tokens': 1024, 'max_num_seqs': 1, 'enable_flashinfer_autotune': False, 'speculative_config': {'method': 'mtp', 'num_speculative_tokens': 1}, 'compilation_config': {'mode': None, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': [], 'ir_enable_torch_wrap': None, 'splitting_ops': None, 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': None, 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': None, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [1, 2], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': None, 'pass_config': {}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': None, 'static_all_moe_layers': []}, 'kernel_config': KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=[], fused_add_rms_norm=[]), enable_flashinfer_autotune=None, enable_cutedsl_warmup=False, enable_jit_warmup=False, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')}
(APIServer pid=59) WARNING 08-30 23:38:57 [envs.py:2187] Unknown vLLM environment variable detected: VLLM_KVQ_TILES
(APIServer pid=59) INFO 08-30 23:38:57 [model.py:679] Resolved architecture: Glm5NextForConditionalGeneration
(APIServer pid=59) INFO 08-30 23:38:57 [model.py:1972] Using max model len 524288
(APIServer pid=59) INFO 08-30 23:39:04 [cache.py:282] Using fp8 data type to store kv cache. It reduces the GPU memory footprint and boosts the performance. Meanwhile, it may cause accuracy drop without a proper scaling factor
(APIServer pid=59) INFO 08-30 23:39:04 [model.py:679] Resolved architecture: Glm5NextMTPModel
(APIServer pid=59) INFO 08-30 23:39:04 [model.py:1972] Using max model len 1048576
(APIServer pid=59) INFO 08-30 23:39:04 [speculative.py:1256] Overriding draft model max model len from 1048576 to 524288
(APIServer pid=59) INFO 08-30 23:39:04 [scheduler.py:242] Chunked prefill is enabled with max_num_batched_tokens=1024.
(APIServer pid=59) INFO 08-30 23:39:04 [config.py:605] Mamba cache mode is set to 'align' for Glm5NextForConditionalGeneration by default when prefix caching is enabled
(APIServer pid=59) WARNING 08-30 23:39:04 [modelopt.py:379] Detected ModelOpt fp8 checkpoint (quant_algo=FP8). Please note that the format is experimental and could change.
(APIServer pid=59) WARNING 08-30 23:39:04 [modelopt.py:1012] Detected ModelOpt NVFP4 checkpoint (quant_algo=NVFP4). Please note that the format is experimental and could change in future.
(APIServer pid=59) WARNING 08-30 23:39:04 [modelopt.py:1012] Detected ModelOpt NVFP4 checkpoint (quant_algo=W4A16_NVFP4). Please note that the format is experimental and could change in future.
(APIServer pid=59) WARNING 08-30 23:39:04 [modelopt.py:1676] Detected ModelOpt MXFP8 checkpoint. Please note that the format is experimental and could change in future.
(APIServer pid=59) INFO 08-30 23:39:04 [vllm.py:1319] Auto-enabling VLLM_USE_BREAKABLE_CUDAGRAPH=1. Set VLLM_USE_BREAKABLE_CUDAGRAPH=0 to opt out.
(APIServer pid=59) INFO 08-30 23:39:04 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
(APIServer pid=59) WARNING 08-30 23:39:04 [vllm.py:1849] max_num_scheduled_tokens is set to 1024 based on the speculative decoding settings. This may lead to suboptimal performance. Consider increasing max_num_batched_tokens to accommodate the additional draft token slots, or decrease num_speculative_tokens.
(APIServer pid=59) INFO 08-30 23:39:04 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
(APIServer pid=59) [transformers] The following generation flags are not valid and may be ignored: ['top_p']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
(EngineCore pid=327) INFO 08-30 23:39:12 [core.py:122] Initializing a V1 LLM engine (v0.1.dev20051+g487ecf187) with config: model='/model', speculative_config=SpeculativeConfig(method='mtp', model='/model', num_spec_tokens=1), tokenizer='/model', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=524288, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=modelopt_mixed, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='glm45', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=glm-max, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [1024], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.FULL_AND_PIECEWISE: (2, 1)>, 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 2, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=False, enable_cutedsl_warmup=False, enable_jit_warmup=False, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto')
(EngineCore pid=327) INFO 08-30 23:39:12 [multiproc_executor.py:149] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=192.168.1.80 (local), world_size=2, local_world_size=2
(Worker pid=461) INFO 08-30 23:39:20 [parallel_state.py:1638] world_size=2 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_ff8e5afc5b1143acad305c975dd061fc backend=nccl
(Worker pid=612) INFO 08-30 23:39:24 [parallel_state.py:1638] world_size=2 rank=1 local_rank=1 distributed_init_method=file:///tmp/vllm_dist_ff8e5afc5b1143acad305c975dd061fc backend=nccl
(Worker pid=461) INFO 08-30 23:39:24 [pynccl.py:113] vLLM is using nccl==2.30.7
(Worker pid=612) WARNING 08-30 23:39:25 [symm_mem.py:67] SymmMemCommunicator: Device capability 12.0 not supported, communicator is not available.
(Worker pid=461) WARNING 08-30 23:39:25 [symm_mem.py:67] SymmMemCommunicator: Device capability 12.0 not supported, communicator is not available.
(Worker pid=461) INFO 08-30 23:39:25 [cuda_communicator.py:266] Using ['CUSTOM', 'PYNCCL'] all-reduce backends (in dispatch order) for group 'tp:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL'].
(Worker pid=461) INFO 08-30 23:39:25 [cuda_communicator.py:266] Using ['PYNCCL'] all-reduce backends (in dispatch order) for group 'ep:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL'].
(Worker pid=461) INFO 08-30 23:39:25 [parallel_state.py:1982] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
(Worker pid=461) INFO 08-30 23:39:25 [gpu_worker.py:386] Using V2 Model Runner
(Worker_TP0 pid=461) INFO 08-30 23:39:25 [model_runner.py:353] Loading model from scratch...
(Worker_TP0 pid=461) INFO 08-30 23:39:25 [__init__.py:662] Selected DeepGemmFp8BlockScaledMMKernel for Fp8LinearMethod
(Worker_TP0 pid=461) INFO 08-30 23:39:25 [deep_gemm.py:191] deep_gemm not found in site-packages, trying vendored vllm.third_party.deep_gemm
(Worker_TP0 pid=461) INFO 08-30 23:39:25 [deep_gemm.py:218] DeepGEMM PDL enabled on vllm.third_party.deep_gemm.
(Worker_TP0 pid=461) INFO 08-30 23:39:25 [deep_gemm.py:136] DeepGEMM E8M0 enabled on current platform.
(Worker_TP0 pid=461) INFO 08-30 23:39:25 [cuda.py:595] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention
(Worker_TP0 pid=461) INFO 08-30 23:39:25 [mm_encoder_attention.py:375] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention.
(Worker_TP1 pid=612) INFO 08-30 23:39:26 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
(Worker_TP0 pid=461) WARNING 08-30 23:39:26 [vllm.py:1849] max_num_scheduled_tokens is set to 1024 based on the speculative decoding settings. This may lead to suboptimal performance. Consider increasing max_num_batched_tokens to accommodate the additional draft token slots, or decrease num_speculative_tokens.
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [__init__.py:662] Selected TritonFp8BlockScaledMMKernel for Fp8LinearMethod
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [cuda.py:536] Using FLASHINFER_MLA_SPARSE_SM120 attention backend out of potential backends: ['FLASHINFER_MLA_SPARSE_SM120'].
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [mla_attention.py:475] Using fp8_ds_mla KV cache format for FLASHINFER_MLA_SPARSE_SM120 backend.
(Worker_TP0 pid=461) WARNING 08-30 23:39:26 [mla_attention.py:551] Sparse MLA impl has no dense-MHA prefill path; using the top-k MQA path only.
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [nvfp4.py:291] Using 'FLASHINFER_CUTLASS' NvFp4 MoE backend out of potential backends: ['FLASHINFER_TRTLLM', 'FLASHINFER_CUTEDSL', 'FLASHINFER_CUTLASS', 'VLLM_CUTLASS', 'MARLIN', 'HUMMING', 'EMULATION'].
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [__init__.py:662] Selected DeepGemmFp8BlockScaledMMKernel for _Fp8BlockLMHeadMethod
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [weight_utils.py:858] Filesystem type for checkpoints: EXT4. Checkpoint size: 173.36 GiB. Available RAM: 396.63 GiB.
(Worker_TP0 pid=461) INFO 08-30 23:39:26 [weight_utils.py:881] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 0% Completed | 0/121 [00:00<?, ?it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 1% Completed | 1/121 [00:00<00:25, 4.64it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 2% Completed | 2/121 [00:00<00:32, 3.72it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 2% Completed | 3/121 [00:01<00:48, 2.45it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 3% Completed | 4/121 [00:01<00:37, 3.09it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 4% Completed | 5/121 [00:01<00:32, 3.61it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 5% Completed | 6/121 [00:01<00:31, 3.66it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 6% Completed | 7/121 [00:02<00:31, 3.60it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 7% Completed | 8/121 [00:02<00:38, 2.95it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 7% Completed | 9/121 [00:02<00:33, 3.33it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 8% Completed | 10/121 [00:03<00:33, 3.31it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 9% Completed | 11/121 [00:03<00:38, 2.83it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 10% Completed | 12/121 [00:03<00:35, 3.06it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 11% Completed | 13/121 [00:04<00:36, 2.93it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 12% Completed | 14/121 [00:04<00:33, 3.19it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 12% Completed | 15/121 [00:04<00:32, 3.31it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 13% Completed | 16/121 [00:05<00:35, 2.93it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 14% Completed | 17/121 [00:05<00:34, 3.03it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 15% Completed | 18/121 [00:05<00:39, 2.61it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 16% Completed | 19/121 [00:06<00:35, 2.88it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 17% Completed | 20/121 [00:06<00:30, 3.28it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 17% Completed | 21/121 [00:06<00:26, 3.71it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 18% Completed | 22/121 [00:06<00:29, 3.36it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 19% Completed | 23/121 [00:07<00:26, 3.68it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 20% Completed | 24/121 [00:07<00:29, 3.32it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 21% Completed | 25/121 [00:07<00:28, 3.32it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 21% Completed | 26/121 [00:08<00:28, 3.33it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 22% Completed | 27/121 [00:08<00:30, 3.12it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 23% Completed | 28/121 [00:08<00:28, 3.24it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 24% Completed | 29/121 [00:09<00:28, 3.26it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 25% Completed | 30/121 [00:09<00:30, 3.01it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 26% Completed | 31/121 [00:09<00:33, 2.73it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 26% Completed | 32/121 [00:10<00:30, 2.96it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 27% Completed | 33/121 [00:10<00:29, 2.99it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 28% Completed | 34/121 [00:10<00:28, 3.07it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 29% Completed | 35/121 [00:11<00:25, 3.40it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 30% Completed | 36/121 [00:11<00:26, 3.21it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 31% Completed | 37/121 [00:11<00:25, 3.28it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 31% Completed | 38/121 [00:11<00:23, 3.48it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 32% Completed | 39/121 [00:12<00:22, 3.67it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 33% Completed | 40/121 [00:12<00:22, 3.64it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 34% Completed | 41/121 [00:12<00:20, 3.98it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 35% Completed | 42/121 [00:12<00:21, 3.74it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 36% Completed | 43/121 [00:13<00:26, 2.95it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 36% Completed | 44/121 [00:13<00:24, 3.15it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 37% Completed | 45/121 [00:13<00:22, 3.45it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 38% Completed | 46/121 [00:14<00:25, 2.95it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 39% Completed | 47/121 [00:14<00:28, 2.61it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 40% Completed | 48/121 [00:15<00:31, 2.31it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 40% Completed | 49/121 [00:15<00:26, 2.71it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 41% Completed | 50/121 [00:15<00:26, 2.72it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 42% Completed | 51/121 [00:16<00:24, 2.81it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 43% Completed | 52/121 [00:16<00:25, 2.70it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 44% Completed | 53/121 [00:17<00:25, 2.62it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 45% Completed | 54/121 [00:17<00:22, 2.93it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 45% Completed | 55/121 [00:17<00:20, 3.28it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 46% Completed | 56/121 [00:17<00:18, 3.51it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 47% Completed | 57/121 [00:18<00:19, 3.34it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 48% Completed | 58/121 [00:18<00:18, 3.47it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 49% Completed | 59/121 [00:18<00:20, 3.04it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 50% Completed | 60/121 [00:19<00:18, 3.27it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 50% Completed | 61/121 [00:19<00:18, 3.29it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 51% Completed | 62/121 [00:19<00:15, 3.72it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 52% Completed | 63/121 [00:19<00:14, 4.07it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 53% Completed | 64/121 [00:19<00:13, 4.32it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 54% Completed | 65/121 [00:20<00:14, 3.97it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 55% Completed | 66/121 [00:20<00:12, 4.40it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 55% Completed | 67/121 [00:20<00:10, 4.91it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 56% Completed | 68/121 [00:20<00:10, 5.23it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 57% Completed | 69/121 [00:20<00:10, 4.98it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 58% Completed | 70/121 [00:21<00:09, 5.12it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 59% Completed | 71/121 [00:21<00:12, 4.09it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 60% Completed | 72/121 [00:21<00:11, 4.28it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 60% Completed | 73/121 [00:21<00:10, 4.46it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 61% Completed | 74/121 [00:22<00:10, 4.49it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 62% Completed | 75/121 [00:22<00:10, 4.59it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 63% Completed | 76/121 [00:22<00:10, 4.29it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 64% Completed | 77/121 [00:22<00:10, 4.34it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 64% Completed | 78/121 [00:23<00:09, 4.61it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 65% Completed | 79/121 [00:23<00:09, 4.65it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 66% Completed | 80/121 [00:23<00:08, 4.75it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 67% Completed | 81/121 [00:23<00:08, 4.58it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 68% Completed | 82/121 [00:23<00:09, 4.33it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 69% Completed | 83/121 [00:24<00:08, 4.59it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 69% Completed | 84/121 [00:24<00:07, 4.91it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 70% Completed | 85/121 [00:24<00:07, 4.53it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 71% Completed | 86/121 [00:24<00:08, 4.11it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 72% Completed | 87/121 [00:25<00:08, 3.87it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 73% Completed | 88/121 [00:25<00:08, 3.95it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 74% Completed | 89/121 [00:25<00:07, 4.39it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 74% Completed | 90/121 [00:25<00:07, 4.30it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 75% Completed | 91/121 [00:26<00:08, 3.69it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 76% Completed | 92/121 [00:26<00:07, 4.04it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 77% Completed | 93/121 [00:26<00:06, 4.45it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 78% Completed | 94/121 [00:26<00:05, 4.75it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 79% Completed | 95/121 [00:26<00:05, 5.09it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 79% Completed | 96/121 [00:27<00:05, 4.94it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 80% Completed | 97/121 [00:27<00:05, 4.52it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 81% Completed | 98/121 [00:27<00:04, 4.95it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 82% Completed | 99/121 [00:27<00:04, 4.70it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 83% Completed | 100/121 [00:27<00:04, 5.00it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 83% Completed | 101/121 [00:28<00:04, 4.26it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 84% Completed | 102/121 [00:28<00:05, 3.51it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 85% Completed | 103/121 [00:28<00:04, 4.06it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 88% Completed | 106/121 [00:28<00:01, 7.63it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 89% Completed | 108/121 [00:29<00:02, 6.30it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 90% Completed | 109/121 [00:29<00:01, 6.08it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 91% Completed | 110/121 [00:29<00:02, 5.21it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 92% Completed | 111/121 [00:30<00:02, 4.61it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 93% Completed | 112/121 [00:30<00:01, 4.89it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 93% Completed | 113/121 [00:30<00:01, 4.64it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 94% Completed | 114/121 [00:30<00:01, 4.93it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 95% Completed | 115/121 [00:30<00:01, 4.76it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 96% Completed | 116/121 [00:31<00:00, 5.17it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 97% Completed | 117/121 [00:31<00:00, 5.73it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 98% Completed | 118/121 [00:31<00:00, 6.08it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 98% Completed | 119/121 [00:31<00:00, 6.39it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 99% Completed | 120/121 [00:31<00:00, 6.91it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 100% Completed | 121/121 [00:33<00:00, 1.50it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 100% Completed | 121/121 [00:33<00:00, 3.61it/s]
(Worker_TP0 pid=461)
(Worker_TP0 pid=461) INFO 08-30 23:40:00 [default_loader.py:430] Loading weights took 33.52 seconds
[rank1]:[W830 23:40:02.195190924 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 635437056 bytes (free: 14155776, total: 101975851008).
(Worker_TP0 pid=461) WARNING 08-30 23:40:02 [modelopt.py:1534] w1_weight_scale_2 must match w3_weight_scale_2. Accuracy may be affected.
(Worker_TP0 pid=461) INFO 08-30 23:40:02 [nvfp4.py:564] Using MoEPrepareAndFinalizeNoDPEPModular
[rank0]:[W830 23:40:02.661052083 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 635437056 bytes (free: 14155776, total: 101975851008).
(Worker_TP0 pid=461) INFO 08-30 23:40:02 [weight_utils.py:858] Filesystem type for checkpoints: EXT4. Checkpoint size: 173.36 GiB. Available RAM: 396.49 GiB.
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 0% Completed | 0/121 [00:00<?, ?it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 7% Completed | 8/121 [00:00<00:01, 72.01it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 13% Completed | 16/121 [00:00<00:01, 70.26it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 20% Completed | 24/121 [00:00<00:01, 69.11it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 26% Completed | 32/121 [00:00<00:01, 70.09it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 33% Completed | 40/121 [00:00<00:01, 69.50it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 39% Completed | 47/121 [00:00<00:01, 69.11it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 45% Completed | 54/121 [00:00<00:00, 68.07it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 50% Completed | 61/121 [00:00<00:00, 68.47it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 56% Completed | 68/121 [00:00<00:00, 68.88it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 62% Completed | 75/121 [00:01<00:00, 68.71it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 68% Completed | 82/121 [00:01<00:00, 67.88it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 74% Completed | 89/121 [00:01<00:00, 67.27it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 79% Completed | 96/121 [00:01<00:00, 66.30it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 85% Completed | 103/121 [00:01<00:00, 65.84it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 91% Completed | 110/121 [00:01<00:00, 35.20it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 97% Completed | 117/121 [00:02<00:00, 41.13it/s]
(Worker_TP0 pid=461) Loading safetensors checkpoint shards: 100% Completed | 121/121 [00:02<00:00, 46.97it/s]
(Worker_TP0 pid=461)
(Worker_TP0 pid=461) INFO 08-30 23:40:05 [default_loader.py:430] Loading weights took 2.58 seconds
(Worker_TP0 pid=461) WARNING 08-30 23:40:05 [speculator.py:191] Draft model Glm5NextMTP does not support external multimodal embeddings. Embeddings from the target model will not be passed to the drafter; using text-only draft inputs instead.
(Worker_TP1 pid=612) INFO 08-30 23:40:05 [model_runner.py:374] Model loading took 87.28 GiB and 40.227077 seconds
(Worker_TP1 pid=612) INFO 08-30 23:40:05 [interface.py:635] Setting kv cache block size to 64 for DEEPSEEK_V32_INDEXER backend.
(Worker_TP1 pid=612) INFO 08-30 23:40:05 [interface.py:926] Setting attention block size to 5120 tokens to ensure that attention page size is >= mamba page size.
(Worker_TP1 pid=612) INFO 08-30 23:40:05 [interface.py:950] Padding mamba page size by 3.54% to ensure that mamba page size and attention page size are exactly equal.
(Worker_TP0 pid=461) INFO 08-30 23:40:06 [model_runner.py:374] Model loading took 87.28 GiB and 40.743755 seconds
(Worker_TP0 pid=461) INFO 08-30 23:40:06 [topk_topp_sampler.py:62] Using FlashInfer for top-p & top-k sampling.
(Worker_TP0 pid=461) INFO 08-30 23:40:06 [interface.py:635] Setting kv cache block size to 64 for DEEPSEEK_V32_INDEXER backend.
(Worker_TP0 pid=461) INFO 08-30 23:40:06 [interface.py:926] Setting attention block size to 5120 tokens to ensure that attention page size is >= mamba page size.
(Worker_TP0 pid=461) INFO 08-30 23:40:06 [interface.py:950] Padding mamba page size by 3.54% to ensure that mamba page size and attention page size are exactly equal.
(EngineCore pid=327) INFO 08-30 23:40:06 [torch_utils.py:262] Reducing Torch threads from 24 to 1 for serving. Set OMP_NUM_THREADS in the external environment to override.
(Worker_TP0 pid=461) INFO 08-30 23:40:07 [encoder_runner.py:119] Encoder cache will be initialized with a budget of 7921 tokens, and profiled with 1 image items of the maximum feature size.
(Worker_TP1 pid=612) 2026-08-30 23:40:16 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
(Worker_TP1 pid=612) 2026-08-30 23:40:20 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
(Worker_TP1 pid=612) WARNING 08-30 23:40:21 [rocm.py:42] Failed to import from amdsmi with ModuleNotFoundError("No module named 'amdsmi'")
(Worker_TP1 pid=612) WARNING 08-30 23:40:21 [rocm.py:47] Failed to import from vllm._C with ModuleNotFoundError("No module named 'vllm._C'")
(Worker_TP1 pid=612) WARNING 08-30 23:40:21 [rocm.py:58] Failed to import from vllm._rocm_C with ModuleNotFoundError("No module named 'vllm._rocm_C'")
(Worker_TP1 pid=612) INFO 08-30 23:40:21 [fp8_utils.py:843] Using configuration from /usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/quantization/utils/configs/N=12576,K=4096,device_name=NVIDIA_RTX_PRO_6000_Blackwell_Workstation_Edition,dtype=fp8_w8a8,block_shape=[32,32].json for W8A8 Block FP8 kernel.
(Worker_TP0 pid=461) 2026-08-30 23:40:16 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
(Worker_TP0 pid=461) 2026-08-30 23:40:20 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
(Worker_TP0 pid=461) WARNING 08-30 23:40:21 [rocm.py:42] Failed to import from amdsmi with ModuleNotFoundError("No module named 'amdsmi'")
(Worker_TP0 pid=461) WARNING 08-30 23:40:21 [rocm.py:47] Failed to import from vllm._C with ModuleNotFoundError("No module named 'vllm._C'")
(Worker_TP0 pid=461) WARNING 08-30 23:40:21 [rocm.py:58] Failed to import from vllm._rocm_C with ModuleNotFoundError("No module named 'vllm._rocm_C'")
(Worker_TP0 pid=461) INFO 08-30 23:40:21 [fp8_utils.py:843] Using configuration from /usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/quantization/utils/configs/N=12576,K=4096,device_name=NVIDIA_RTX_PRO_6000_Blackwell_Workstation_Edition,dtype=fp8_w8a8,block_shape=[32,32].json for W8A8 Block FP8 kernel.
(Worker_TP1 pid=612) /usr/local/lib/python3.12/dist-packages/triton/language/core.py:2284: UserWarning: tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.
(Worker_TP1 pid=612) warn("tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.")
(Worker_TP0 pid=461) /usr/local/lib/python3.12/dist-packages/triton/language/core.py:2284: UserWarning: tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.
(Worker_TP0 pid=461) warn("tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.")
[23:40:21] : Warning: T.vectorized loop over `i_hci` with extent 4 is lowered as a serial loop because TileLang could not find a valid vectorization plan. Scalar accumulator updates inside the loop are a common cause; move reductions to T.unroll or T.serial if this is intended.
[23:40:21] : Warning: T.vectorized loop over `i_hci` with extent 4 is lowered as a serial loop because TileLang could not find a valid vectorization plan. Scalar accumulator updates inside the loop are a common cause; move reductions to T.unroll or T.serial if this is intended.
(Worker_TP0 pid=461) 2026-08-30 23:40:21 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_post_tilelang` with `out_idx=None`
(Worker_TP0 pid=461) 2026-08-30 23:40:22 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_post_tilelang`
(Worker_TP0 pid=461) INFO 08-30 23:40:25 [gpu_worker.py:492] Initial free memory 93.92 GiB, reserved 3.24 GiB memory for KV Cache as specified by kv_cache_memory_bytes config and skipped memory profiling. This does not respect the gpu_memory_utilization config. Only use kv_cache_memory_bytes config when you want manual control of KV cache memory size. If OOM'ed, check the difference of initial free memory between the current run and the previous run where kv_cache_memory_bytes is suggested and update it correspondingly.
(Worker_TP1 pid=612) 2026-08-30 23:40:21 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_post_tilelang` with `out_idx=None`
(Worker_TP1 pid=612) 2026-08-30 23:40:22 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_post_tilelang`
(Worker_TP1 pid=612) INFO 08-30 23:40:25 [gpu_worker.py:492] Initial free memory 93.92 GiB, reserved 3.24 GiB memory for KV Cache as specified by kv_cache_memory_bytes config and skipped memory profiling. This does not respect the gpu_memory_utilization config. Only use kv_cache_memory_bytes config when you want manual control of KV cache memory size. If OOM'ed, check the difference of initial free memory between the current run and the previous run where kv_cache_memory_bytes is suggested and update it correspondingly.
(EngineCore pid=327) INFO 08-30 23:40:25 [kv_cache_utils.py:2141] GPU KV cache size: 547,486 tokens, Maximum concurrency for 524,288 tokens per request: 1.04x
(Worker_TP0 pid=461) INFO 08-30 23:40:25 [indexer.py:762] DSA indexer decode path: use_flattening=False supports_varlen=False (next_n=2, use_fp4_indexer_cache=False)
(Worker_TP0 pid=461) INFO 08-30 23:40:25 [kernel_warmup.py:95] Warming up ll_bf16 router GEMM kernels for shapes: ((4096, 288),).
(Worker_TP0 pid=461) DeepGEMM warmup: 0%| | 0/732 [00:00<?, ?it/s] DeepGEMM warmup: 100%|██████████| 732/732 [00:00<00:00, 16553.87it/s]
(Worker_TP0 pid=461) INFO 08-30 23:40:26 [kernel_warmup.py:170] Skipping FlashInfer autotune because it is disabled.
(Worker_TP0 pid=461) INFO 08-30 23:40:26 [breakable_cudagraph.py:288] Breakable CUDA graph enabled
(Worker_TP0 pid=461) Capturing CUDA graphs (PIECEWISE): 0%| | 0/2 [00:00<?, ?it/s] Capturing CUDA graphs (PIECEWISE): 0%| | 0/2 [00:00<?, ?it/s] Capturing CUDA graphs (PIECEWISE): 0%| | 0/2 [00:04<?, ?it/s](Worker_TP1 pid=612) /usr/local/lib/python3.12/dist-packages/triton/language/core.py:2284: UserWarning: tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.
(Worker_TP1 pid=612) warn("tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.")
/usr/local/lib/python3.12/dist-packages/triton/language/core.py:2284: UserWarning: tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.
(Worker_TP0 pid=461) warn("tl.make_block_ptr is deprecated. Use TensorDescriptor or tl.make_tensor_descriptor instead.")
(Worker_TP0 pid=461) Capturing CUDA graphs (PIECEWISE): 0%| | 0/2 [00:46<?, ?it/s] Capturing CUDA graphs (PIECEWISE): 0%| | 0/2 [00:47<?, ?it/s] Capturing CUDA graphs (PIECEWISE): 0%| | 0/2 [00:47<?, ?it/s] Capturing CUDA graphs (PIECEWISE): 0%| | 0/2 [00:51<?, ?it/s] Capturing CUDA graphs (PIECEWISE): 50%|█████ | 1/2 [00:53<00:53, 53.60s/it] Capturing CUDA graphs (PIECEWISE): 100%|██████████| 2/2 [00:55<00:00, 23.03s/it] Capturing CUDA graphs (PIECEWISE): 100%|██████████| 2/2 [00:55<00:00, 27.61s/it]
(Worker_TP0 pid=461) Capturing CUDA graphs (FULL): 0%| | 0/1 [00:00<?, ?it/s](Worker_TP0 pid=461) 2026-08-30 23:40:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
(Worker_TP0 pid=461) 2026-08-30 23:40:31 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
(Worker_TP0 pid=461) 2026-08-30 23:41:13 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_fused_tilelang` with `out_idx=None`
(Worker_TP0 pid=461) 2026-08-30 23:41:13 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_fused_tilelang`
(Worker_TP0 pid=461) 2026-08-30 23:41:14 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
(Worker_TP0 pid=461) 2026-08-30 23:41:18 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
Capturing CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 2.53it/s] Capturing CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 2.53it/s]
(Worker_TP1 pid=612) 2026-08-30 23:40:26 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
(Worker_TP1 pid=612) 2026-08-30 23:40:31 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
(Worker_TP1 pid=612) 2026-08-30 23:41:13 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_fused_tilelang` with `out_idx=None`
(Worker_TP1 pid=612) 2026-08-30 23:41:13 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_fused_tilelang`
(Worker_TP1 pid=612) 2026-08-30 23:41:14 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
(Worker_TP1 pid=612) 2026-08-30 23:41:18 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
(Worker_TP1 pid=612) INFO 08-30 23:41:22 [speculator.py:148] Capturing model for speculator...
(Worker_TP0 pid=461) INFO 08-30 23:41:22 [speculator.py:148] Capturing model for speculator...
(Worker_TP0 pid=461) Capturing prefill CUDA graphs (PIECEWISE): 0%| | 0/2 [00:00<?, ?it/s] Capturing prefill CUDA graphs (PIECEWISE): 50%|█████ | 1/2 [00:00<00:00, 1.51it/s] Capturing prefill CUDA graphs (PIECEWISE): 100%|██████████| 2/2 [00:01<00:00, 1.60it/s] Capturing prefill CUDA graphs (PIECEWISE): 100%|██████████| 2/2 [00:01<00:00, 1.58it/s]
(Worker_TP0 pid=461) Capturing prefill CUDA graphs (FULL): 0%| | 0/1 [00:00<?, ?it/s] Capturing prefill CUDA graphs (FULL): 100%|██████████| 1/1 [00:00<00:00, 111.29it/s]
(Worker_TP1 pid=612) INFO 08-30 23:41:23 [model_runner.py:1012] Graph capturing finished in 57 secs, took 0.61 GiB
(Worker_TP0 pid=461) INFO 08-30 23:41:23 [model_runner.py:1012] Graph capturing finished in 57 secs, took 0.61 GiB
(Worker_TP0 pid=461) INFO 08-30 23:41:26 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
(Worker_TP1 pid=612) INFO 08-30 23:41:26 [jit_monitor.py:79] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
(EngineCore pid=327) INFO 08-30 23:41:26 [shm_broadcast.py:801] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization).
(Worker_TP0 pid=461) INFO 08-30 23:41:26 [torch_utils.py:262] Reducing Torch threads from 24 to 1 for serving. Set OMP_NUM_THREADS in the external environment to override.
(EngineCore pid=327) INFO 08-30 23:41:26 [core.py:363] init engine (profile, create kv cache, warmup model) took 80.44 s
(EngineCore pid=327) INFO 08-30 23:41:30 [kernel.py:306] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
(EngineCore pid=327) WARNING 08-30 23:41:30 [vllm.py:1849] max_num_scheduled_tokens is set to 1024 based on the speculative decoding settings. This may lead to suboptimal performance. Consider increasing max_num_batched_tokens to accommodate the additional draft token slots, or decrease num_speculative_tokens.
(EngineCore pid=327) INFO 08-30 23:41:30 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant
(APIServer pid=59) INFO 08-30 23:41:31 [api_server.py:678] Supported tasks: ['generate']
(APIServer pid=59) INFO 08-30 23:41:31 [parser_manager.py:37] "auto" tool choice has been enabled.
(APIServer pid=59) INFO 08-30 23:41:32 [hf.py:551] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this.
(APIServer pid=59) INFO 08-30 23:41:34 [base.py:235] Multi-modal warmup completed in 1.661s
(APIServer pid=59) INFO 08-30 23:41:36 [base.py:235] Readonly multi-modal warmup completed in 1.725s
(APIServer pid=59) WARNING 08-30 23:41:36 [model.py:1720] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 1.0, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
(APIServer pid=59) INFO 08-30 23:41:39 [api_server.py:682] Starting vLLM server on http://0.0.0.0:18998
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:51] Available routes are:
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /openapi.json, Methods: HEAD, GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /docs, Methods: HEAD, GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /docs/oauth2-redirect, Methods: HEAD, GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /redoc, Methods: HEAD, GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /load, Methods: GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /version, Methods: GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /health, Methods: GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /metrics, Methods: GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /tokenize, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /detokenize, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/models, Methods: GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /ping, Methods: GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /ping, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /invocations, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/chat/completions, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/chat/completions/batch, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/responses, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/responses/{response_id}, Methods: GET
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/responses/{response_id}/cancel, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/completions, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/messages, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/messages/count_tokens, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /generative_scoring, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /scale_elastic_ep, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /is_scaling_elastic_ep, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/chat/completions/render, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/completions/render, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/chat/completions/derender, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /v1/completions/derender, Methods: POST
(APIServer pid=59) INFO 08-30 23:41:39 [launcher.py:60] Route: /inference/v1/generate, Methods: POST
(APIServer pid=59) INFO: Started server process [59]
(APIServer pid=59) INFO: Waiting for application startup.
(APIServer pid=59) INFO: Application startup complete.
(APIServer pid=59) INFO: 127.0.0.1:56012 - "GET /health HTTP/1.1" 200 OK
(Worker_TP1 pid=612) INFO 08-30 23:40:05 [model_runner.py:374] Model loading took 87.28 GiB and 40.227077 seconds
(Worker_TP0 pid=461) INFO 08-30 23:40:06 [model_runner.py:374] Model loading took 87.28 GiB and 40.743755 seconds
(EngineCore pid=327) INFO 08-30 23:40:25 [kv_cache_utils.py:2141] GPU KV cache size: 547,486 tokens, Maximum concurrency for 524,288 tokens per request: 1.04x
(Worker_TP0 pid=461) WARNING 08-30 23:41:41 [jit_monitor.py:135] Triton kernel JIT compilation during inference: BuildPrefillChunkMetadataKernel.kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
(Worker_TP0 pid=461) WARNING 08-30 23:41:41 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _w8a8_triton_block_scaled_mm. This causes a latency spike; consider extending warmup to cover this shape/config.
(Worker_TP0 pid=461) WARNING 08-30 23:41:42 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _kpool_softmax_rotate_write_cache_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
(Worker_TP0 pid=461) WARNING 08-30 23:41:42 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _compute_local_logits_stats_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
(Worker_TP0 pid=461) WARNING 08-30 23:41:42 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _rejection_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
(Worker_TP0 pid=461) WARNING 08-30 23:41:42 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _resample_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
(APIServer pid=59) INFO: 127.0.0.1:56028 - "POST /v1/chat/completions HTTP/1.1" 200 OK
TEXT: PASS (391)
(APIServer pid=59) INFO: 127.0.0.1:56040 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(Worker_TP0 pid=461) WARNING 08-30 23:41:43 [jit_monitor.py:135] Triton kernel JIT compilation during inference: rotary_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
(Worker_TP0 pid=461) WARNING 08-30 23:41:43 [jit_monitor.py:135] TileLang JIT compilation during inference: mhc_pre_big_fuse_with_norm_tilelang. This causes a latency spike; consider extending warmup to cover this shape/config.
(APIServer pid=59) INFO: 127.0.0.1:56054 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:56070 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:56074 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO 08-30 23:41:49 [loggers.py:310] Engine 000: Avg prompt throughput: 34.4 tokens/s, Avg generation throughput: 6.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 6.8%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=59) INFO 08-30 23:41:49 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.96, Accepted throughput: 2.94 tokens/s, Drafted throughput: 3.05 tokens/s, Accepted: 55 tokens, Drafted: 57 tokens, Per-position acceptance rate: 0.965, Avg Draft acceptance rate: 96.5%
(APIServer pid=59) INFO: 127.0.0.1:56078 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:56090 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:56096 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:56108 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:56118 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:56124 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51220 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(Worker_TP0 pid=461) 2026-08-30 23:41:43 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:129): TileLang begins to compile kernel `mhc_pre_big_fuse_with_norm_tilelang` with `out_idx=None`
(Worker_TP0 pid=461) 2026-08-30 23:41:48 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang completes to compile kernel `mhc_pre_big_fuse_with_norm_tilelang`
(Worker_TP0 pid=461) WARNING 08-30 23:41:51 [jit_monitor.py:135] Triton kernel JIT compilation during inference: _kpool_tail_seed_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
(APIServer pid=59) INFO: 127.0.0.1:51222 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51238 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51244 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51256 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51266 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51272 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51278 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51284 - "POST /v1/chat/completions HTTP/1.1" 200 OK
[img01.png] OK img_toks=258 ans='Red'
[img02.png] OK img_toks=258 ans='SATURN'
[img03.png] OK img_toks=258 ans='4'
[img04.png] OK img_toks=258 ans='Triangle'
[img05.png] OK img_toks=258 ans='7318'
[img06.png] OK img_toks=258 ans='3'
[img07.png] OK img_toks=258 ans='Circle'
[img08.png] OK img_toks=258 ans='MANGO'
[img09.png] OK img_toks=258 ans='5'
[img10.png] OK img_toks=258 ans='The black square is on the **left** side of the image.'
VISION VERDICT: PASS (10/10 correct, gate >=10, hard_fail=False)
wrote results-vision-512kconf-8666.json
=== 1x ~501k fresh prefill ===
(APIServer pid=59) INFO: 127.0.0.1:51296 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51304 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51318 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51332 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51340 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO 08-30 23:41:59 [loggers.py:310] Engine 000: Avg prompt throughput: 249.9 tokens/s, Avg generation throughput: 61.3 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
(APIServer pid=59) INFO 08-30 23:41:59 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.89, Accepted throughput: 28.20 tokens/s, Drafted throughput: 31.79 tokens/s, Accepted: 283 tokens, Drafted: 319 tokens, Per-position acceptance rate: 0.887, Avg Draft acceptance rate: 88.7%
(APIServer pid=59) INFO: 127.0.0.1:51346 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51354 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51358 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:51366 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:46572 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:46580 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:46582 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:46592 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:46608 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:46616 - "POST /tokenize HTTP/1.1" 200 OK
(APIServer pid=59) INFO 08-30 23:42:09 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 8.5%, Prefix cache hit rate: 0.0%, MM cache hit rate: 0.0%
[rank1]:[W830 23:42:34.181939325 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 178257920 bytes (free: 114819072, total: 101975851008).
[rank0]:[W830 23:42:34.182028155 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 178257920 bytes (free: 127401984, total: 101975851008).
[rank1]:[W830 23:42:38.152985005 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 203423744 bytes (free: 158859264, total: 101975851008).
[rank0]:[W830 23:42:38.490482835 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 205520896 bytes (free: 68681728, total: 101975851008).
[rank1]:[W830 23:42:41.840248434 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 226492416 bytes (free: 217579520, total: 101975851008).
[rank0]:[W830 23:42:42.181948344 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 228589568 bytes (free: 211288064, total: 101975851008).
[rank1]:[W830 23:42:45.224976732 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 247463936 bytes (free: 211288064, total: 101975851008).
[rank0]:[W830 23:42:45.571221181 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 249561088 bytes (free: 207093760, total: 101975851008).
[rank1]:[W830 23:42:48.301529558 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 266338304 bytes (free: 267911168, total: 101975851008).
[rank0]:[W830 23:42:48.649412232 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 268435456 bytes (free: 265814016, total: 101975851008).
[rank1]:[W830 23:42:51.403194296 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 285212672 bytes (free: 98041856, total: 101975851008).
[rank0]:[W830 23:42:51.753405922 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 287309824 bytes (free: 95944704, total: 101975851008).
[rank1]:[W830 23:42:54.183804863 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 301989888 bytes (free: 230162432, total: 101975851008).
[rank0]:[W830 23:42:54.536068703 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 304087040 bytes (free: 230162432, total: 101975851008).
[rank1]:[W830 23:42:56.986131102 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 318767104 bytes (free: 95944704, total: 101975851008).
[rank0]:[W830 23:42:57.340623246 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 320864256 bytes (free: 95944704, total: 101975851008).
[rank1]:[W830 23:42:59.449005812 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 333447168 bytes (free: 295174144, total: 101975851008).
[rank0]:[W830 23:42:59.805219431 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 335544320 bytes (free: 297271296, total: 101975851008).
[rank1]:[W830 23:43:01.931529503 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 348127232 bytes (free: 192413696, total: 101975851008).
[rank0]:[W830 23:43:02.290645974 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 350224384 bytes (free: 194510848, total: 101975851008).
[rank1]:[W830 23:43:04.427191813 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 362807296 bytes (free: 89653248, total: 101975851008).
[rank0]:[W830 23:43:04.788776567 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 364904448 bytes (free: 91750400, total: 101975851008).
[rank1]:[W830 23:43:06.582338268 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 375390208 bytes (free: 362283008, total: 101975851008).
[rank0]:[W830 23:43:06.945909107 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 377487360 bytes (free: 366477312, total: 101975851008).
[rank1]:[W830 23:43:08.749475760 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 387973120 bytes (free: 286785536, total: 101975851008).
[rank0]:[W830 23:43:09.114796003 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 390070272 bytes (free: 290979840, total: 101975851008).
[rank1]:[W830 23:43:10.928961237 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 400556032 bytes (free: 211288064, total: 101975851008).
[rank0]:[W830 23:43:11.298048332 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 402653184 bytes (free: 215482368, total: 101975851008).
[rank1]:[W830 23:43:13.120925840 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 413138944 bytes (free: 135790592, total: 101975851008).
[rank0]:[W830 23:43:13.491285390 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 415236096 bytes (free: 139984896, total: 101975851008).
[rank1]:[W830 23:43:15.324590440 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 425721856 bytes (free: 60293120, total: 101975851008).
[rank0]:[W830 23:43:15.695674317 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 427819008 bytes (free: 64487424, total: 101975851008).
[rank1]:[W830 23:43:17.172276925 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 436207616 bytes (free: 421003264, total: 101975851008).
[rank0]:[W830 23:43:17.545006047 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 438304768 bytes (free: 427294720, total: 101975851008).
[rank1]:[W830 23:43:19.028108396 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 446693376 bytes (free: 368574464, total: 101975851008).
[rank0]:[W830 23:43:19.401915915 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 448790528 bytes (free: 374865920, total: 101975851008).
[rank1]:[W830 23:43:20.896195487 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 457179136 bytes (free: 316145664, total: 101975851008).
[rank0]:[W830 23:43:21.270943542 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 459276288 bytes (free: 322437120, total: 101975851008).
[rank1]:[W830 23:43:22.740643999 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 467664896 bytes (free: 263716864, total: 101975851008).
[rank0]:[W830 23:43:23.108960854 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 469762048 bytes (free: 270008320, total: 101975851008).
[rank1]:[W830 23:43:24.564337548 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 478150656 bytes (free: 211288064, total: 101975851008).
[rank0]:[W830 23:43:24.932404841 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 480247808 bytes (free: 217579520, total: 101975851008).
[rank1]:[W830 23:43:26.388998346 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 488636416 bytes (free: 158859264, total: 101975851008).
[rank0]:[W830 23:43:26.759605073 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 490733568 bytes (free: 165150720, total: 101975851008).
[rank1]:[W830 23:43:28.217846118 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 499122176 bytes (free: 106430464, total: 101975851008).
[rank0]:[W830 23:43:28.590325529 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 501219328 bytes (free: 112721920, total: 101975851008).
[rank1]:[W830 23:43:30.057423902 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 1 while trying to allocate 509607936 bytes (free: 54001664, total: 101975851008).
[rank0]:[W830 23:43:30.428232634 CUDACachingAllocator.cpp:3933] memory allocation failed with OOM on device 0 while trying to allocate 511705088 bytes (free: 60293120, total: 101975851008).
(APIServer pid=59) INFO: 127.0.0.1:46622 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:50372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:50386 - "POST /v1/chat/completions HTTP/1.1" 200 OK
calibration: 0.18182 tok/char on 9471 chars
[round 1] built prompt: 434531 tokens (2390199 chars)
[round 1] extended to 500974 tokens
[round 1] prompt_tokens=501006 completion=55 wall=84.7s finish=stop ok=True
[round 1] reply: 'ACKNOWLEDGED. The text appears to be random filler content with no meaningful message.'
[round 1] decode at 501006 ctx: 61 tok in 3.0s (1-tok call 2.7s) -> ~272.1 tok/s
PASS 1/1
=== post-500k vision re-probe ===
(APIServer pid=59) INFO: 127.0.0.1:50388 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:50398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO 08-30 23:43:39 [loggers.py:310] Engine 000: Avg prompt throughput: 52082.9 tokens/s, Avg generation throughput: 40.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 6.8%, Prefix cache hit rate: 65.2%, MM cache hit rate: 0.0%
(APIServer pid=59) INFO 08-30 23:43:39 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 1.96, Accepted throughput: 1.94 tokens/s, Drafted throughput: 2.03 tokens/s, Accepted: 194 tokens, Drafted: 203 tokens, Per-position acceptance rate: 0.956, Avg Draft acceptance rate: 95.6%
(APIServer pid=59) INFO: 127.0.0.1:50408 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:50420 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:48808 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:48818 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:48832 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:48848 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:48850 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:48858 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:48874 - "POST /v1/chat/completions HTTP/1.1" 200 OK
(APIServer pid=59) INFO: 127.0.0.1:48882 - "POST /v1/chat/completions HTTP/1.1" 200 OK
[PASS] shapes_3red_circles: gold='3' got='3' (1.0s)
[PASS] shapes_5blue_squares: gold='5' got='5' (1.1s)
[PASS] shapes_2green_circles: gold='2' got='2' (0.8s)
[PASS] shapes_4yellow_squares: gold='4' got='4' (1.2s)
[PASS] text_mango: gold='MANGO' got='MANGO' (0.3s)
[PASS] text_river7: gold='RIVER 7' got='RIVER 7' (0.3s)
[PASS] text_zebra_code: gold='ZK-42' got='ZK-42' (0.4s)
[PASS] text_plum: gold='PLUM BASKET' got='PLUM BASKET' (0.4s)
[PASS] chart_tallest_mar: gold='MAR' got='MAR' (0.4s)
[PASS] chart_shortest_q2: gold='Q2' got='Q2' (0.4s)
[PASS] chart_value_feb: gold='75' got='75' (0.4s)
[PASS] chart_pie_cats: gold='CATS' got='**CATS**' (0.9s)
SCORE 512kconf-post500k-8666: 12/12
wrote /home/ian/h3-v2/glm-max/results-vision-512kconf-post500k-8666.json
=== min free watermark (MiB): 110 ===
VISFP8-512K done
[rewarm] posting completion to prod glm-5.3-flash via 8080...
[rewarm] answer: 391 17 × 23 = 17 × 23
Let me compute: 17 × 23 = 17 × 20 + 17 × 3 = 340 + 51 = 391
[rewarm] running: glm-5.3-flash=ready
[rewarm] PROD WARM OK
[bash] released