clm-0136
The Unsloth Qwen3.8-Flash-Next UD-Q4_K_XL quant served by Gufo (gufo-org/gufo 1b4e682, HIP) with its shared-Q8 MTP head, single session at 200192 context behind the slotpin proxy, is the fleet's production serving config on aibeast from 2026-10-03 (cfg-0188). ADMITTED: (1) tau2 airline tasks 0-25 = 22/26 (0.846, run-0674) at reasoning_effort=low, matched to the same quant on the Unsloth llama.cpp Vulkan fork (24/26, run-0632), N=1 each, with all four failures traced to model decisions (two omitted required figures, one transfer to a human, one cancel-and-rebook) and none to the engine's tool path; (2) 4.47 Wh per correct answer (eng-0281) vs 9.90 for the same quant on llama.cpp (eng-0275), reasoning-matched; (3) a served tokenizer-exact code depth sweep holds 42-58 t/s decode and ~1300-1420 t/s prefill from 32768 to 196608 tokens (run-0675..run-0678); (4) live tau2 workload decode median 48.2 t/s, warm TTFT median 0.66 s. BLOCKED: no claim that Gufo is more capable than llama.cpp on this quant (22 vs 24 of 26 is within single-trial variation and was not repeated); no cross-quant ranking against the ROCmFP4-FAST-v2 series (cfg-0182, default reasoning); the depth sweep is N=1 per cell and code-content only.