CUDA error: an illegal memory access was encountered using vllm(sm120)
(APIServer pid=16130) INFO: 127.0.0.1:46668 - "POST /v1/completions HTTP/1.1" 200 OK
(EngineCore pid=16399) 2026-07-01 11:39:03,655 - WARNING - autotuner.py:1147 - flashinfer.jit: [AutoTuner]: No tuned config covers fp8_gemm input_shapes=(torch.Size([1, 3060, 5120]), torch.Size([1, 5120, 16384]), torch.Size([]), torch.Size([]), torch.Size([1, 3060, 16384]), torch.Size([33554688])); falling back to runner=CutlassFp8GemmRunner tactic=-1. This shape is outside the tuning bucket range -- expand tuning_buckets / max_num_tokens during the next tuning pass to avoid this perf cliff.
(APIServer pid=16130) INFO 07-01 11:39:06 [loggers.py:273] Engine 000: Avg prompt throughput: 4434.3 tokens/s, Avg generation throughput: 5.6 tokens/s, Running: 0 reqs, Waiting: 976 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(EngineCore pid=16399) 2026-07-01 11:39:13,493 - WARNING - autotuner.py:1147 - flashinfer.jit: [AutoTuner]: No tuned config covers fp8_gemm input_shapes=(torch.Size([1, 2548, 6144]), torch.Size([1, 6144, 5120]), torch.Size([]), torch.Size([]), torch.Size([1, 2548, 5120]), torch.Size([33554688])); falling back to runner=CutlassFp8GemmRunner tactic=-1. This shape is outside the tuning bucket range -- expand tuning_buckets / max_num_tokens during the next tuning pass to avoid this perf cliff.
(APIServer pid=16130) INFO 07-01 11:39:16 [loggers.py:273] Engine 000: Avg prompt throughput: 4418.9 tokens/s, Avg generation throughput: 6.8 tokens/s, Running: 0 reqs, Waiting: 908 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=16130) INFO 07-01 11:39:26 [loggers.py:273] Engine 000: Avg prompt throughput: 4266.1 tokens/s, Avg generation throughput: 7.2 tokens/s, Running: 0 reqs, Waiting: 836 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=16130) INFO: 127.0.0.1:46648 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=16130) INFO 07-01 11:39:36 [loggers.py:273] Engine 000: Avg prompt throughput: 4499.3 tokens/s, Avg generation throughput: 8.0 tokens/s, Running: 0 reqs, Waiting: 756 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=16130) INFO 07-01 11:39:46 [loggers.py:273] Engine 000: Avg prompt throughput: 4128.2 tokens/s, Avg generation throughput: 7.6 tokens/s, Running: 0 reqs, Waiting: 936 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=16130) INFO 07-01 11:39:56 [loggers.py:273] Engine 000: Avg prompt throughput: 4402.5 tokens/s, Avg generation throughput: 8.4 tokens/s, Running: 0 reqs, Waiting: 852 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(EngineCore pid=16399) ERROR 07-01 11:39:57 [dump_input.py:72] Dumping input data for V1 LLM engine (v0.24.0) with config: model='nvidia/Qwen3.6-27B-NVFP4', speculative_config=SpeculativeConfig(method='mtp', model='nvidia/Qwen3.6-27B-NVFP4', num_spec_tokens=3), tokenizer='nvidia/Qwen3.6-27B-NVFP4', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=262144, download_dir=None, load_format=fastsafetensors, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=modelopt_mixed, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=fp8, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='qwen3', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_verbose=False), seed=0, served_model_name=nvidia/Qwen3.6-27B-NVFP4, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.VLLM_COMPILE: 3>, 'debug_dump_path': None, 'cache_dir': '/home/admin2/.cache/vllm/torch_compile_cache/50a7361981', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::linear_attention', 'vllm::plamo2_mamba_mixer', 'vllm::qwen_gdn_attention_core', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::kda_attention', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [8192], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.PIECEWISE: 1>, 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8, 16, 24, 32], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 32, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': '/home/admin2/.cache/vllm/torch_compile_cache/50a7361981/rank_0_0/eagle_head', 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, moe_backend='marlin', linear_backend='auto'),
(EngineCore pid=16399) ERROR 07-01 11:39:57 [dump_input.py:79] Dumping scheduler output for model execution: SchedulerOutput(scheduled_new_reqs=[], scheduled_cached_reqs=CachedRequestData(req_ids=[],resumed_req_ids=set(),new_token_ids_lens=[],all_token_ids_lens={},new_block_ids=[],num_computed_tokens=[],num_output_tokens=[]), num_scheduled_tokens={}, total_num_scheduled_tokens=0, scheduled_spec_decode_tokens={}, scheduled_encoder_inputs={}, num_common_prefix_blocks=[0, 0, 0, 0], finished_req_ids=[], free_encoder_mm_hashes=[], preempted_req_ids=[], has_structured_output_requests=false, pending_structured_output_tokens=false, num_invalid_spec_tokens=null, kv_connector_metadata=null, ec_connector_metadata=null, new_block_ids_to_zero=null, num_spec_tokens_to_schedule=3)
terminate called after throwing an instance of 'c10::AcceleratorError'
(EngineCore pid=16399) ERROR 07-01 11:39:57 [dump_input.py:81] Dumping scheduler stats: SchedulerStats(num_running_reqs=4, num_waiting_reqs=840, num_skipped_waiting_reqs=0, step_counter=0, current_wave=0, kv_cache_usage=0.18055555555555558, prefix_cache_stats=PrefixCacheStats(reset=False, requests=0, queries=0, hits=0, preempted_requests=0, preempted_queries=0, preempted_hits=0), connector_prefix_cache_stats=None, kv_cache_eviction_events=[], spec_decoding_stats=None, kv_connector_stats=None, waiting_lora_adapters={}, running_lora_adapters={}, cudagraph_stats=None, perf_stats=None)
what(): CUDA error: an illegal memory access was encountered
Search for cudaErrorIllegalAddress' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information. CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1 Compile with TORCH_USE_CUDA_DSA` to enable device-side assertions.
Exception raised from currentStreamCaptureStatusMayInitCtx at /pytorch/c10/cuda/CUDAGraphsC10Utils.h:71 (most recent call first):
frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits, std::allocator >) + 0x9d (0x74afa797305d in /home/admin2/.vllm/lib/python3.12/site-packages/torch/lib/libc10.so)
frame #1: + 0xc114 (0x74afa7d00114 in /home/admin2/.vllm/lib/python3.12/site-packages/torch/lib/libc10_cuda.so)
frame #2: + 0xcc775e (0x74aef30c775e in /home/admin2/.vllm/lib/python3.12/site-packages/torch/lib/libtorch_cuda.so)
frame #3: + 0x7f8e4 (0x74afa79548e4 in /home/admin2/.vllm/lib/python3.12/site-packages/torch/lib/libc10.so)
frame #4: c10::TensorImpl::~TensorImpl() + 0x9 (0x74afa794e279 in /home/admin2/.vllm/lib/python3.12/site-packages/torch/lib/libc10.so)
frame #5: + 0x869d55 (0x74af20269d55 in /home/admin2/.vllm/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
frame #6: + 0x869df1 (0x74af20269df1 in /home/admin2/.vllm/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
frame #7: VLLM::EngineCore() [0x59f2e3]
frame #8: + 0x14adad (0x74aef174adad in /home/admin2/.vllm/lib/python3.12/site-packages/numpy/_core/_multiarray_umath.cpython-312-x86_64-linux-gnu.so)
frame #9: + 0x14adad (0x74aef174adad in /home/admin2/.vllm/lib/python3.12/site-packages/numpy/_core/_multiarray_umath.cpython-312-x86_64-linux-gnu.so)
frame #10: VLLM::EngineCore() [0x59178a]
frame #11: VLLM::EngineCore() [0x59f2e3]
frame #12: + 0xa004 (0x74af2f7a4004 in /home/admin2/.vllm/lib/python3.12/site-packages/msgspec/_core.cpython-312-x86_64-linux-gnu.so)
frame #13: VLLM::EngineCore() [0x55bbf4]
frame #14: + 0xa004 (0x74af2f7a4004 in /home/admin2/.vllm/lib/python3.12/site-packages/msgspec/_core.cpython-312-x86_64-linux-gnu.so)
frame #15: _PyEval_EvalFrameDefault + 0x9a6b (0x5e036b in VLLM::EngineCore)
frame #16: VLLM::EngineCore() [0x54cded]
frame #17: _PyEval_EvalFrameDefault + 0x4c32 (0x5db532 in VLLM::EngineCore)
frame #18: VLLM::EngineCore() [0x54ce52]
frame #19: VLLM::EngineCore() [0x6f858c]
frame #20: VLLM::EngineCore() [0x6b954c]
frame #21: + 0x9caa4 (0x74afa969caa4 in /lib/x86_64-linux-gnu/libc.so.6)
frame #22: + 0x129c6c (0x74afa9729c6c in /lib/x86_64-linux-gnu/libc.so.6)
(APIServer pid=16130) INFO 07-01 11:40:06 [loggers.py:273] Engine 000: Avg prompt throughput: 411.2 tokens/s, Avg generation throughput: 0.8 tokens/s, Running: 0 reqs, Waiting: 844 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=16130) INFO 07-01 11:40:16 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 844 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0%
(APIServer pid=16130) WARNING 07-01 11:41:05 [core_client.py:702] [shutdown] MPClient: engine core exited unexpectedly; starting cleanup
(APIServer pid=16130) INFO 07-01 11:41:05 [core_client.py:655] [shutdown] MPClient: start timeout=default
(APIServer pid=16130) INFO 07-01 11:41:05 [core_client.py:657] [shutdown] MPClient: stopping engine manager
(APIServer pid=16130) INFO 07-01 11:41:05 [core_client.py:659] [shutdown] MPClient: engine manager stopped
(APIServer pid=16130) INFO 07-01 11:41:05 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
(APIServer pid=16130) INFO 07-01 11:41:05 [core_client.py:662] [shutdown] MPClient: complete
(APIServer pid=16130) ERROR 07-01 11:41:05 [async_llm.py:704] AsyncLLM output_handler failed.
(APIServer pid=16130) ERROR 07-01 11:41:05 [async_llm.py:704] Traceback (most recent call last):
(APIServer pid=16130) ERROR 07-01 11:41:05 [async_llm.py:704] File "/home/admin2/.vllm/lib/python3.12/site-packages/vllm/v1/engine/async_llm.py", line 660, in output_handler
(APIServer pid=16130) ERROR 07-01 11:41:05 [async_llm.py:704] outputs = await engine_core.get_output_async()
(APIServer pid=16130) ERROR 07-01 11:41:05 [async_llm.py:704] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
(APIServer pid=16130) ERROR 07-01 11:41:05 [async_llm.py:704] File "/home/admin2/.vllm/lib/python3.12/site-packages/vllm/v1/engine/core_client.py", line 1061, in get_output_async
(APIServer pid=16130) ERROR 07-01 11:41:05 [async_llm.py:704] raise self._format_exception(outputs) from None
(APIServer pid=16130) ERROR 07-01 11:41:05 [async_llm.py:704] vllm.v1.engine.exceptions.EngineDeadError: EngineCore encountered an issue. See stack trace (above) for the root cause.
(APIServer pid=16130) INFO: 127.0.0.1:46660 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46666 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46668 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46648 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46648 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46660 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46668 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46666 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46648 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46660 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46668 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46666 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46648 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46660 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46668 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46666 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46648 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46660 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46668 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46666 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46648 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46660 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46668 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46666 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46648 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: Shutting down
(APIServer pid=16130) INFO: 127.0.0.1:46660 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46668 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
(APIServer pid=16130) INFO: 127.0.0.1:46666 - "POST /v1/completions HTTP/1.1" 500 Internal Server Error
A restart usually works out fine for me