Home › Evidence › Records › clm-0136

clm-0136

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0136/ — this address never moves; the anchor /records/#clm-0136 keeps resolving

The Unsloth Qwen3.8-Flash-Next UD-Q4_K_XL quant served by Gufo (gufo-org/gufo 1b4e682, HIP) with its shared-Q8 MTP head, single session at 200192 context behind the slotpin proxy, is the fleet's production serving config on aibeast from 2026-10-03 (cfg-0188). ADMITTED: (1) tau2 airline tasks 0-25 = 22/26 (0.846, run-0674) at reasoning_effort=low, matched to the same quant on the Unsloth llama.cpp Vulkan fork (24/26, run-0632), N=1 each, with all four failures traced to model decisions (two omitted required figures, one transfer to a human, one cancel-and-rebook) and none to the engine's tool path; (2) 4.47 Wh per correct answer (eng-0281) vs 9.90 for the same quant on llama.cpp (eng-0275), reasoning-matched; (3) a served tokenizer-exact code depth sweep holds 42-58 t/s decode and ~1300-1420 t/s prefill from 32768 to 196608 tokens (run-0675..run-0678); (4) live tau2 workload decode median 48.2 t/s, warm TTFT median 0.66 s. BLOCKED: no claim that Gufo is more capable than llama.cpp on this quant (22 vs 24 of 26 is within single-trial variation and was not repeated); no cross-quant ranking against the ROCmFP4-FAST-v2 series (cfg-0182, default reasoning); the depth sweep is N=1 per cell and code-content only.

verified 2026-10-03 · volatility high
evidence cfg-0188 run-0674 eng-0281 run-0675 run-0676 run-0677 run-0678 run-0632 eng-0275

Note — the record's own working

Gufo main moves daily; the measured commit is pinned in cfg-0188. Rollback to cfg-0182 is a two-service swap.

Cited by — computed at build time, never stored

model pages qwen38-flash-next
candidate gate history qwen38-flash-next