Home › Evidence › Records › clm-0054

clm-0054

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0054/ — this address never moves; the anchor /records/#clm-0054 keeps resolving

On gfx1151 at f16 KV, Qwen3.8-27B (dense) is a fourth model measured under the per-model, per-phase backend rule (clm-0050), and it lands emphatically on the prefill-favours-ROCm side: DECODE is backend-independent on this model (every matched cell within 5%), but Vulkan PREFILL collapses with depth — 0.55x ROCm at d32768 on Q8_0 (87.82 vs 159.98 t/s) and 0.48x on UD-Q4_K_XL (99.80 vs 207.61) — and at d131072 stock Vulkan cannot complete the cell at all: 2 of 2 reps in BOTH quants aborted with vk::DeviceLostError, the kernel logging an amdgpu ring timeout and recovering the device by ring reset, while ROCm completed every cell it was offered (83.44 pp / 6.11 tg Q8_0, 95.59 pp / 8.14 tg UD-Q4_K_XL at d131072). This REVERSES the screen's own shallow-data backend pick: a decode-only d0 probe chose Vulkan (+0.8% Q8_0, +5.0% UD-Q4_K_XL), and depth then showed prefill swinging to ROCm by +82.2% (Q8_0) and +108.0% (UD-Q4_K_XL) at d32768 and Vulkan failing outright where ROCm ran. For this model on this chip, ROCm is the only backend that holds to 131k, and the d0 pick was wrong in the way clm-0050 predicts shallow picks to be wrong: backend advantage is per-model AND per-phase, and it moves with depth.

Note — the record's own working

METHOD — qwen38-screen phase B: stock llama.cpp 3653e6d, one build per backend arm, f16 KV, fa on, -ngl 999, --load-mode none (clm-0053's rule, carried by the queue from its first rep), pp1024/tg256, llama-bench defaults otherwise (-b 2048 / -ub 512), median of 3 fresh-process reps with page cache dropped between reps. Depth-major order, d131072 last. Widest rep spread in any completed cell: 3.08% pp (vulkan Q8_0 d32768); every other cell under 1%. pp1024 / tg256 by depth: | quant | depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---|---| | Q8_0 | 0 | 226.22 / 7.81 | 249.63 / 7.86 | 1.103 | 1.006 | | Q8_0 | 32,768 | 159.98 / 7.29 | 87.82 / 7.27 | 0.549 | 0.998 | | Q8_0 | 131,072 | 83.44 / 6.11 | DEVICE LOST | — | — | | UD-Q4_K_XL | 0 | 346.49 / 11.44 | 361.69 / 12.03 | 1.044 | 1.052 | | UD-Q4_K_XL | 32,768 | 207.61 / 10.36 | 99.80 / 10.69 | 0.481 | 1.031 | | UD-Q4_K_XL | 131,072 | 95.59 / 8.14 | DEVICE LOST | — | — | THE DEVICE-LOST CELLS are recorded as failed runs with the process and kernel evidence verbatim (run-0265, run-0266), the same discipline as the OOM cells of the Nemotron matrix (run-0234/run-0235): a cell that failed is a result, and no invented number stands in for it. The failure is recoverable — the kernel reset the compute ring each time ("Ring comp_1.2.0 reset succeeded ... device wedged, but recovered through reset") and the box needed no reboot — but four out of four attempted reps across two quants at d131072 is deterministic enough for the queue's fast-fail rule. This is the first in-lab reproduction of the failure class llama.cpp issue 25664 documents for Vulkan/RADV at deep context on this hardware class, which clm-0050 has carried as a reliability counterweight since 2026-08-13: the Qwen3.6-35B matrix ran Vulkan clean to 131k, this model does not. SCOPE — the d204800 cell was not run in this screen (the queue's depth list ended at 131072); whether ROCm's prefill lead extends there is measured on other models but only extrapolated here. The decode parity is this model's OWN result: on Qwen3.6-35B Vulkan leads decode by 19-21% at every depth, on gpt-oss-120b by 10-16%, on this model by at most 5.2% at d0 and effectively nothing at depth — another way the rule is per-model. The backend verdict for serving this model at useful context is ROCm on both phases: prefill by a wide margin, decode by indifference, and completion-at-depth by necessity. The community's first-throughput rows for this model (r/StrixHalo, Vulkan fork b10283, amd_iommu=off, pp512/tg128, 2 reps) are adjacent evidence, not the same measurement — different build lineage, IOMMU state, and phase lengths; their own A/B puts amd_iommu=off alone at up to +34-38% on dense prefill. No public ROCm figures for this model on this silicon existed before this matrix.

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b