ANNOUNCE box benchbox · cards 0+1 (both RTX 3090s, 250 W caps) · gap 5 · vLLM 0.29.0 serve openjev/openjev-FP8 TP=2 with --limit-mm-per-prompt image:1 · start 2026-09-22T10:00:57Z · expect ~10 min to ready, then ~10 min of task arms (APIServer pid=7671) INFO 09-22 10:01:05 [api_utils.py:347] (APIServer pid=7671) INFO 09-22 10:01:05 [api_utils.py:347] █ █ █▄ ▄█ (APIServer pid=7671) INFO 09-22 10:01:05 [api_utils.py:347] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.29.0 (APIServer pid=7671) INFO 09-22 10:01:05 [api_utils.py:347] █▄█▀ █ █ █ █ model /workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10 (APIServer pid=7671) INFO 09-22 10:01:05 [api_utils.py:347] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀ (APIServer pid=7671) INFO 09-22 10:01:05 [api_utils.py:347] (APIServer pid=7671) INFO 09-22 10:01:05 [api_utils.py:286] non-default args: {'model_tag': '/workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10', 'host': '127.0.0.1', 'model': '/workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10', 'trust_remote_code': True, 'max_model_len': 16384, 'quantization': 'fp8', 'max_logprobs': 64, 'served_model_name': ['openjev-fp8'], 'tensor_parallel_size': 2, 'gpu_memory_utilization': 0.9, 'enable_prefix_caching': True, 'limit_mm_per_prompt': {'image': 1}, 'max_num_seqs': 1, 'gdn_prefill_backend': 'triton'} (APIServer pid=7671) INFO 09-22 10:01:05 [model.py:684] Resolved architecture: Qwen3_5ForConditionalGeneration (APIServer pid=7671) INFO 09-22 10:01:05 [model.py:2021] Using max model len 16384 (APIServer pid=7671) INFO 09-22 10:01:07 [config.py:625] Mamba cache mode is set to 'align' for Qwen3_5ForConditionalGeneration by default when prefix caching is enabled (APIServer pid=7671) INFO 09-22 10:01:07 [kernel.py:369] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']) (APIServer pid=7671) INFO 09-22 10:01:07 [compilation.py:329] Enabled custom fusions: norm_quant, act_quant (APIServer pid=7671) [transformers] The `use_fast` parameter is deprecated and will be removed in a future version. Use `backend="torchvision"` instead of `use_fast=True`, or `backend="pil"` instead of `use_fast=False`. (EngineCore pid=7839) INFO 09-22 10:01:21 [core.py:123] Initializing a V1 LLM engine (v0.29.0) with config: model='/workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10', speculative_config=None, tokenizer='/workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=16384, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=fp8, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=openjev-fp8, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['+quant_fp8', 'none', '+quant_fp8'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::qwen4_exp_compute_ple_ngram_ids', 'vllm::qwen4_exp_ple_short_conv', 'vllm::qwen4_exp_qsa_with_output', 'vllm::linear_attention', 'vllm::qwen_gdn_attention_core', 'vllm::qwen_gdn_attention_core_fused_norm_packed', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::hpc_rope_norm_forward', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 2, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto') (EngineCore pid=7839) INFO 09-22 10:01:21 [multiproc_executor.py:153] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip= (local), world_size=2, local_world_size=2 (Worker pid=7948) INFO 09-22 10:01:29 [parallel_state.py:1775] world_size=2 rank=0 local_rank=0 distributed_init_method=file:///tmp/vllm_dist_66ada92dc13042e4a112676c6a7adc5e backend=nccl (Worker pid=7949) INFO 09-22 10:01:29 [parallel_state.py:1775] world_size=2 rank=1 local_rank=1 distributed_init_method=file:///tmp/vllm_dist_66ada92dc13042e4a112676c6a7adc5e backend=nccl (Worker pid=7948) INFO 09-22 10:01:30 [pynccl.py:113] vLLM is using nccl==2.29.7 (Worker pid=7948) WARNING 09-22 10:01:30 [symm_mem.py:67] SymmMemCommunicator: Device capability 8.6 not supported, communicator is not available. (Worker pid=7949) WARNING 09-22 10:01:30 [symm_mem.py:67] SymmMemCommunicator: Device capability 8.6 not supported, communicator is not available. (Worker pid=7949) WARNING 09-22 10:01:30 [flashinfer_all_reduce.py:383] FlashInfer All Reduce is disabled because it is not supported for world_size=2. (Worker pid=7948) WARNING 09-22 10:01:30 [flashinfer_all_reduce.py:383] FlashInfer All Reduce is disabled because it is not supported for world_size=2. (Worker pid=7949) WARNING 09-22 10:01:30 [custom_all_reduce.py:252] Custom allreduce is disabled because your platform lacks GPU P2P capability or P2P test failed. To silence this warning, specify disable_custom_all_reduce=True explicitly. (Worker pid=7948) WARNING 09-22 10:01:30 [custom_all_reduce.py:252] Custom allreduce is disabled because your platform lacks GPU P2P capability or P2P test failed. To silence this warning, specify disable_custom_all_reduce=True explicitly. (Worker pid=7948) INFO 09-22 10:01:30 [cuda_communicator.py:269] Using ['PYNCCL'] all-reduce backends (in dispatch order) for group 'tp:0' out of potential backends: ['FLASHINFER', 'NCCL_SYMM_MEM', 'QUICK_REDUCE', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL']. (Worker pid=7948) INFO 09-22 10:01:30 [parallel_state.py:2119] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A (Worker pid=7948) INFO 09-22 10:01:30 [gpu_worker.py:429] Using V2 Model Runner (Worker_TP0 pid=7948) INFO 09-22 10:01:30 [model_runner.py:382] Loading model from scratch... (Worker_TP0 pid=7948) INFO 09-22 10:01:30 [cuda.py:551] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention (Worker_TP0 pid=7948) INFO 09-22 10:01:30 [mm_encoder_attention.py:372] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention. (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] Module vllm.third_party.deep_gemm was found but failed to import (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] Traceback (most recent call last): (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "/workshop/bench-vllm-venv/lib/python3.12/site-packages/vllm/utils/import_utils.py", line 406, in _has_module (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] importlib.import_module(module_name) (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "/workshop/.local/share/uv/python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/importlib/__init__.py", line 90, in import_module (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] return _bootstrap._gcd_import(name[level:], package, level) (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 1387, in _gcd_import (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 1360, in _find_and_load (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 1331, in _find_and_load_unlocked (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 935, in _load_unlocked (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 999, in exec_module (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 488, in _call_with_frames_removed (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "/workshop/bench-vllm-venv/lib/python3.12/site-packages/vllm/third_party/deep_gemm/__init__.py", line 126, in (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] _find_cuda_home() # CUDA home (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] ^^^^^^^^^^^^^^^^^ (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] File "/workshop/bench-vllm-venv/lib/python3.12/site-packages/vllm/third_party/deep_gemm/__init__.py", line 120, in _find_cuda_home (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] assert cuda_home is not None (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP0 pid=7948) WARNING 09-22 10:01:30 [import_utils.py:408] AssertionError (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] Module vllm.third_party.deep_gemm was found but failed to import (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] Traceback (most recent call last): (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "/workshop/bench-vllm-venv/lib/python3.12/site-packages/vllm/utils/import_utils.py", line 406, in _has_module (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] importlib.import_module(module_name) (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "/workshop/.local/share/uv/python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/importlib/__init__.py", line 90, in import_module (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] return _bootstrap._gcd_import(name[level:], package, level) (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 1387, in _gcd_import (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 1360, in _find_and_load (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 1331, in _find_and_load_unlocked (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 935, in _load_unlocked (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 999, in exec_module (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "", line 488, in _call_with_frames_removed (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "/workshop/bench-vllm-venv/lib/python3.12/site-packages/vllm/third_party/deep_gemm/__init__.py", line 126, in (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] _find_cuda_home() # CUDA home (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] ^^^^^^^^^^^^^^^^^ (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] File "/workshop/bench-vllm-venv/lib/python3.12/site-packages/vllm/third_party/deep_gemm/__init__.py", line 120, in _find_cuda_home (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] assert cuda_home is not None (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] ^^^^^^^^^^^^^^^^^^^^^ (Worker_TP1 pid=7949) WARNING 09-22 10:01:30 [import_utils.py:408] AssertionError (Worker_TP0 pid=7948) INFO 09-22 10:01:30 [__init__.py:695] Selected MarlinFP8ScaledMMLinearKernel for Fp8LinearMethod (Worker_TP0 pid=7948) INFO 09-22 10:01:30 [qwen_gdn_linear_attn.py:167] Using Triton/FLA GDN prefill kernel (requested=triton, head_k_dim=128). (Worker_TP0 pid=7948) INFO 09-22 10:01:30 [qwen_gdn_linear_attn.py:519] GDN decode kernel: cuda (Worker_TP0 pid=7948) INFO 09-22 10:01:30 [cuda.py:492] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. (Worker_TP0 pid=7948) INFO 09-22 10:01:30 [flash_attn.py:897] Using FlashAttention version 2 (Worker_TP0 pid=7948) INFO 09-22 10:01:31 [weight_utils.py:863] Filesystem type for checkpoints: EXT4. Checkpoint size: 28.30 GiB. Available RAM: 23.34 GiB. (Worker_TP0 pid=7948) INFO 09-22 10:01:31 [weight_utils.py:893] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre) and the checkpoint size (28.30 GiB) exceeds 90% of available RAM (23.34 GiB). (Worker_TP0 pid=7948) Loading safetensors checkpoint shards: 0% Completed | 0/12 [00:00= mamba page size. (Worker_TP1 pid=7949) INFO 09-22 10:01:51 [interface.py:942] Padding mamba page size by 0.13% to ensure that mamba page size and attention page size are exactly equal. (Worker_TP0 pid=7948) Loading safetensors checkpoint shards: 100% Completed | 12/12 [00:20<00:00, 1.33s/it] (Worker_TP0 pid=7948) Loading safetensors checkpoint shards: 100% Completed | 12/12 [00:20<00:00, 1.70s/it] (Worker_TP0 pid=7948) (Worker_TP0 pid=7948) INFO 09-22 10:01:51 [default_loader.py:430] Loading weights took 20.45 seconds (Worker_TP0 pid=7948) WARNING 09-22 10:01:51 [marlin_utils_fp8.py:112] Your GPU does not have native support for FP8 computation but FP8 quantization is being used. Weight-only FP8 compression will be used leveraging the Marlin kernel. This may degrade performance for compute-heavy workloads. (Worker_TP0 pid=7948) INFO 09-22 10:01:52 [model_runner.py:404] Model loading took 14.73 GiB memory and 21.758062 seconds (Worker_TP0 pid=7948) INFO 09-22 10:01:52 [topk_topp_sampler.py:46] FlashInfer top-p/top-k sampling disabled via VLLM_USE_FLASHINFER_SAMPLER=0. (Worker_TP0 pid=7948) INFO 09-22 10:01:52 [interface.py:918] Setting attention block size to 784 tokens to ensure that attention page size is >= mamba page size. (Worker_TP0 pid=7948) INFO 09-22 10:01:52 [interface.py:942] Padding mamba page size by 0.13% to ensure that mamba page size and attention page size are exactly equal. (EngineCore pid=7839) INFO 09-22 10:01:52 [torch_utils.py:277] Reducing Torch threads from 8 to 1 for serving. Set OMP_NUM_THREADS in the external environment to override. (EngineCore pid=7839) INFO 09-22 10:01:52 [utils.py:306] Using LBNHC KV cache layout. (Worker_TP1 pid=7949) [transformers] The `use_fast` parameter is deprecated and will be removed in a future version. Use `backend="torchvision"` instead of `use_fast=True`, or `backend="pil"` instead of `use_fast=False`. (Worker_TP0 pid=7948) [transformers] The `use_fast` parameter is deprecated and will be removed in a future version. Use `backend="torchvision"` instead of `use_fast=True`, or `backend="pil"` instead of `use_fast=False`. (Worker_TP0 pid=7948) INFO 09-22 10:01:59 [encoder_runner.py:131] Encoder cache will be initialized with a budget of 16384 tokens, and profiled with 1 image items of the maximum feature size. (Worker_TP0 pid=7948) INFO 09-22 10:02:13 [caching.py:343] reconstructed serializable fn from standalone compile artifacts. num_artifacts=21 num_submods=65 (Worker_TP1 pid=7949) INFO 09-22 10:02:13 [caching.py:343] reconstructed serializable fn from standalone compile artifacts. num_artifacts=21 num_submods=65 (Worker_TP0 pid=7948) INFO 09-22 10:02:13 [decorators.py:313] Directly load AOT compilation from path /workshop/.cache/vllm/torch_compile_cache/torch_aot_compile/ad843f4aad913fd62d3b0045cdfa29cbf9ef52233b455cd095ab86f268bf4562/rank_0_0/model (Worker_TP0 pid=7948) INFO 09-22 10:02:13 [monitor.py:53] torch.compile took 0.82 s in total (Worker_TP1 pid=7949) INFO 09-22 10:02:13 [decorators.py:313] Directly load AOT compilation from path /workshop/.cache/vllm/torch_compile_cache/torch_aot_compile/ad843f4aad913fd62d3b0045cdfa29cbf9ef52233b455cd095ab86f268bf4562/rank_1_0/model (Worker_TP0 pid=7948) INFO 09-22 10:02:16 [monitor.py:81] Initial profiling/warmup run took 2.28 s (Worker_TP0 pid=7948) Capturing CUDA graphs (PIECEWISE): 0%| | 0/2 [00:00