[2026-09-22T04:09:14Z] ARM3 openjev-fp8-largecard starting — box inferencebox, card largecard GPU-25bc3288-319, one card, no link [2026-09-22T04:09:14Z] snapshot ok: /workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10 (12 shards, rev 4ec320f2) [2026-09-22T04:09:14Z] pre-flight ok: largecard 966 MiB used, no compute app but the reranker [2026-09-22T04:09:14Z] starting vLLM 0.29.0 on 127.0.0.1:8000 — log /workshop/bench-arm3/jev/receipts/vllm-serve-arm3-fp8.log [2026-09-22T04:13:07Z] server answered /v1/models after 233s [2026-09-22T04:13:07Z] ---- the kernel and backend lines, verbatim ---- (APIServer pid=331487) INFO 2026-09-22T04:09:23Z [api_utils.py:286] non-default args: {'model_tag': '/workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10', 'host': '127.0.0.1', 'model': '/workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10', 'trust_remote_code': True, 'max_model_len': 16384, 'quantization': 'fp8', 'max_logprobs': 64, 'served_model_name': ['openjev-fp8-largecard'], 'gpu_memory_utilization': 0.9, 'enable_prefix_caching': True, 'limit_mm_per_prompt': {'image': 0}, 'max_num_seqs': 1, 'gdn_prefill_backend': 'triton'} (EngineCore pid=332027) INFO 2026-09-22T04:09:45Z [core.py:123] Initializing a V1 LLM engine (v0.29.0) with config: model='/workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10', speculative_config=None, tokenizer='/workshop/hf-cache/hub/models--openjev--openjev-FP8/snapshots/4ec320f267401e67c9be04d5df1be4d2b6b64f10', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=16384, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=fp8, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, per_request_spec_decode_metrics='none', kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=openjev-fp8-largecard, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['+quant_fp8', 'none', '+quant_fp8'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::qwen4_exp_compute_ple_ngram_ids', 'vllm::qwen4_exp_ple_short_conv', 'vllm::qwen4_exp_qsa_with_output', 'vllm::linear_attention', 'vllm::qwen_gdn_attention_core', 'vllm::qwen_gdn_attention_core_fused_norm_packed', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::hpc_rope_norm_forward', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [8192], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'enable_qk_norm_rope_fusion': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False, 'fuse_qk_norm_rope_kvcache': False}, 'max_cudagraph_capture_size': 2, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, enable_jit_warmup=True, enable_bf16x3_router_gemm=False, moe_backend='auto', linear_backend='auto') (EngineCore pid=332027) INFO 2026-09-22T04:09:49Z [__init__.py:695] Selected CutlassFp8BlockScaledMMKernel for Fp8LinearMethod (EngineCore pid=332027) INFO 2026-09-22T04:09:49Z [qwen_gdn_linear_attn.py:167] Using Triton/FLA GDN prefill kernel (requested=triton, head_k_dim=128). (EngineCore pid=332027) INFO 2026-09-22T04:09:49Z [qwen_gdn_linear_attn.py:519] GDN decode kernel: cuda (EngineCore pid=332027) INFO 2026-09-22T04:09:49Z [cuda.py:492] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. (EngineCore pid=332027) INFO 2026-09-22T04:09:49Z [flash_attn.py:897] Using FlashAttention version 2 (EngineCore pid=332027) INFO 2026-09-22T04:12:34Z [kernel_warmup.py:309] Using FlashInfer autotune cache file: /workshop/.cache/vllm/flashinfer_autotune_cache/0.6.18/120f/1fd9ddada29197096e86e8102fde69aa7acc27376807256b31e5c0b92ffc9ed1/autotune_configs.json [2026-09-22T04:13:07Z] ---- end kernel lines ---- [2026-09-22T04:13:07Z] running openjev-fp8-largecard: tasks c, a, b — log /workshop/bench-arm3/jev/receipts/run-arm3-fp8.log on-39 - -> top 0.14s on-40 - -> REFUSED: overlong — (18796 tokens) on-41 - -> what-the-adapter-actually-did 0.14s on-42 ['no template tail — no closing sentence of the shape "the article does not specify what was left out, moved or not measured" where the article carries no such statement', "an operator, never the definite one — the hub's voice law (an operator 2026-09-03, re-stated 2026-09-04)", "the prediction's own subject kept — the page predicts that a few COMPANIES will sell intelligence as a utility; the commodity sentence is about SOLD INTELLIGENCE and may not be re-subjected onto the prediction"] -> the-indifferent-prediction 0.13s {"n": 42, "n_reached": 35, "n_labelled": 1, "refused": 7, "refused_by_reason": {"overlong": 7}, "correct": 0, "errors": 0, "brier": 0.5041096348930292, "reliability": [{"lo": 0.0, "hi": 0.1, "n": 0, "sum_p": 0.0, "correct": 0, "mean_p": null, "accuracy": null, "gap": null}, {"lo": 0.1, "hi": 0.2, "n": 0, "sum_p": 0.0, "correct": 0, "mean_p": null, "accuracy": null, "gap": null}, {"lo": 0.2, "hi": 0.3, "n": 0, "sum_p": 0.0, "correct": 0, "mean_p": null, "accuracy": null, "gap": null}, {"lo": 0.3, "hi": 0.4, "n": 0, "sum_p": 0.0, "correct": 0, "mean_p": null, "accuracy": null, "gap": null}, {"lo [2026-09-22T04:14:38Z] RECEIPT arm=openjev-fp8-largecard rc=0 boot=233s task_c=93.5% (101/108) gate_floor=91.5% rows=/workshop/bench-arm3/jev/rows [2026-09-22T04:14:38Z] stopping the bench server (pid 331487) [2026-09-22T04:14:40Z] largecard after the stop: 966 MiB used