qwen38-flash-next-gufo
live
citable URL: https://halobench.com/records/qwen38-flash-next-gufo/ — this address never moves; the anchor /records/#qwen38-flash-next-gufo keeps resolving
Qwen3.8-Flash-Next @ Unsloth UD-Q4_K_XL on Gufo, single session + MTP · kind model · engine igpu · tier daily-driver
runs on aibeast · config cfg-0188
⌁ current state live · ⌁ days in production 8
Lifecycle — append-only
The production serving model on aibeast since 2026-10-03: the Unsloth UD-Q4_K_XL quant with its shared-Q8 MTP head, served by Gufo (gufo-org/gufo, a C++/HIP engine for this APU) as a single session at 200192 context behind the slotpin proxy (gufo-flashnext.service). Same quant as the reasoning=low tau2 baseline, so capability and energy compare like for like: 22/26 on tau2 airline (run-0674) against 24/26 on llama.cpp (run-0632), N=1 each, at 4.47 against 9.90 Wh per correct answer. Served decode holds 42-58 t/s on code from 32K to 196K tokens with prefill near 1300-1400 t/s, and the live tau2 workload ran at a median 48 t/s decode with a 0.66 s warm TTFT. It also leaves about 32 GB of host memory free, against about 14 GB for the previous config. Adopted on serving speed and energy; see clm-0136 for the bounded reading.