<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>HALOBENCH — log</title>
    <link>https://halobench.com/log/</link>
    <atom:link href="https://halobench.com/feed.xml" rel="self" type="application/rss+xml" />
    <description>The lab notebook, generated from the record: claims landed, gate changes, run series, incidents and contributions — retractions badged inline.</description>
    <language>en-gb</language>
    <lastBuildDate>Sat, 15 Aug 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>[claim] On gfx1151 at f16 KV, Qwen3.8-27B (dense) is a fourth model measured under the per-model, per-phase backend rule (clm-0050), and it lands emphatically on the prefill-favours-ROCm side: DECODE is backend-independent on this model (every matched cell within 5%), but Vulkan PREFILL collapses with depth — 0.55x ROCm at d32768 on Q8_0 (87.82 vs 159.98 t/s) and 0.48x on UD-Q4_K_XL (99.80 vs 207.61) — and at d131072 stock Vulkan cannot complete the cell at all: 2 of 2 reps in BOTH quants aborted with vk::DeviceLostError, the kernel logging an amdgpu ring timeout and recovering the device by ring reset, while ROCm completed every cell it was offered (83.44 pp / 6.11 tg Q8_0, 95.59 pp / 8.14 tg UD-Q4_K_XL at d131072).</title>
      <link>https://halobench.com/log/#e-2026-08-15-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-1</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>On gfx1151 at f16 KV, Qwen3.8-27B (dense) is a fourth model measured under the per-model, per-phase backend rule (clm-0050), and it lands emphatically on the prefill-favours-ROCm side: DECODE is backend-independent on this model (every matched cell within 5%), but Vulkan PREFILL collapses with depth — 0.55x ROCm at d32768 on Q8_0 (87.82 vs 159.98 t/s) and 0.48x on UD-Q4_K_XL (99.80 vs 207.61) — and at d131072 stock Vulkan cannot complete the cell at all: 2 of 2 reps in BOTH quants aborted with vk::DeviceLostError, the kernel logging an amdgpu ring timeout and recovering the device by ring reset, while ROCm completed every cell it was offered (83.44 pp / 6.11 tg Q8_0, 95.59 pp / 8.14 tg UD-Q4_K_XL at d131072). — records: clm-0054</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] draft-mtp speculation on Qwen3.8-27B (stock 3653e6d, gfx1151) peaks at spec-draft-n-max=3 over CLEAN cells — Vulkan Q8_0 17.75 t/s (2.26x its 7.86 no-speculation floor, acceptance 0.626), Vulkan UD-Q4_K_XL 26.68 t/s (2.23x, 0.623), ROCm Q8_0 18.25 t/s (2.33x, 0.6235) — and at n_max &gt;= 4 the feature is BROKEN on this model: after accumulated generation volume in a live session (sequential VARIED prompts at full length; not fresh servers, not one repeated prompt, not short generations), generations start terminating at 1 token with &lt;|im_end|&gt; (id 248046).</title>
      <link>https://halobench.com/log/#e-2026-08-15-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-2</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>draft-mtp speculation on Qwen3.8-27B (stock 3653e6d, gfx1151) peaks at spec-draft-n-max=3 over CLEAN cells — Vulkan Q8_0 17.75 t/s (2.26x its 7.86 no-speculation floor, acceptance 0.626), Vulkan UD-Q4_K_XL 26.68 t/s (2.23x, 0.623), ROCm Q8_0 18.25 t/s (2.33x, 0.6235) — and at n_max &gt;= 4 the feature is BROKEN on this model: after accumulated generation volume in a live session (sequential VARIED prompts at full length; not fresh servers, not one repeated prompt, not short generations), generations start terminating at 1 token with &lt;|im_end|&gt; (id 248046). — records: clm-0055</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Qwen3.8-27B's KV cache on llama.cpp costs exactly 64.00 KiB per token — the hybrid-attention allocation working as designed, measured byte-exact on this box: only 16 of the 64 layers carry a conventional KV cache (the 3:1 linear-to-full-attention layout), and 16 layers x 4 KV heads x 256 head_dim x 2 (K+V) x 2 bytes (f16) = 65,536 bytes/token.</title>
      <link>https://halobench.com/log/#e-2026-08-15-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-3</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>Qwen3.8-27B's KV cache on llama.cpp costs exactly 64.00 KiB per token — the hybrid-attention allocation working as designed, measured byte-exact on this box: only 16 of the 64 layers carry a conventional KV cache (the 3:1 linear-to-full-attention layout), and 16 layers x 4 KV heads x 256 head_dim x 2 (K+V) x 2 bytes (f16) = 65,536 bytes/token. — records: clm-0056</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Qwen3.8-27B met its registered agentic prediction: on the standing 5-task tau2 airline smoke subset it scored 1.000 (5/5, 21 tool-call messages, 0 empty assistant turns, VALID under the smoke gate) against the &gt;= 0.80 bar recorded in the candidate record on 2026-08-14, BEFORE any measurement — and against the incumbent qwen36-27b-mtp's 0.80 on the identical tasks under the identical pinned-simulator protocol, where the incumbent failed task 2 and this model passed it.</title>
      <link>https://halobench.com/log/#e-2026-08-15-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-4</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>Qwen3.8-27B met its registered agentic prediction: on the standing 5-task tau2 airline smoke subset it scored 1.000 (5/5, 21 tool-call messages, 0 empty assistant turns, VALID under the smoke gate) against the &gt;= 0.80 bar recorded in the candidate record on 2026-08-14, BEFORE any measurement — and against the incumbent qwen36-27b-mtp's 0.80 on the identical tasks under the identical pinned-simulator protocol, where the incumbent failed task 2 and this model passed it. — records: clm-0057</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] qwen38-27b: listed → screened — Screen tier passed on release day, on-box.</title>
      <link>https://halobench.com/log/#e-2026-08-15-5</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-5</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>qwen38-27b: listed → screened — Screen tier passed on release day, on-box. — records: clm-0055, clm-0056, clm-0057, run-0267</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] qwen38-27b: screened → benched — DAY-ONE SCREEN COMPLETE, all four phases, overnight on release day.</title>
      <link>https://halobench.com/log/#e-2026-08-15-6</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-6</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>qwen38-27b: screened → benched — DAY-ONE SCREEN COMPLETE, all four phases, overnight on release day. — records: clm-0054, clm-0055, clm-0056, clm-0057, run-0271, run-0267, run-0268</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 1 run landed on cfg-0066 (eos-cliffguard)</title>
      <link>https://halobench.com/log/#e-2026-08-15-7</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-7</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>1 run landed on cfg-0066 (eos-cliffguard) — records: run-0267</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 1 run landed on cfg-0067 (mtp-probe)</title>
      <link>https://halobench.com/log/#e-2026-08-15-8</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-8</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>1 run landed on cfg-0067 (mtp-probe) — records: run-0268</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 1 run landed on cfg-0066 (tau2-bench-airline)</title>
      <link>https://halobench.com/log/#e-2026-08-15-9</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-15-9</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>1 run landed on cfg-0066 (tau2-bench-airline) — records: run-0271</description>
      <category>runs</category>
    </item>
    <item>
      <title>[claim] AMENDED 2026-08-14 — the rule below is PER-MODEL AND PER-PHASE, not fleet-wide; see the correction history and the amendment note.</title>
      <link>https://halobench.com/log/#e-2026-08-14-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-1</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>AMENDED 2026-08-14 — the rule below is PER-MODEL AND PER-PHASE, not fleet-wide; see the correction history and the amendment note. — records: clm-0050</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] On gfx1151, stock ROCm 3653e6d's own BF16 KV decode is SLOWER than its own F16 KV decode: 29.26 t/s vs 42.35 t/s at 32,768 depth (Qwen3.6-35B-A3B, fa on) — a 31% penalty for switching to the KV type the community recommends for its quality, with prefill roughly unchanged (587.4 vs 595.0). stew675's rdna-boosts branch (commit ed89854) fixes it with a native BF16 flash-attention tile kernel: 47.31 t/s BF16 decode, +61.7% over stock's own BF16 and stock's best config on this chip.</title>
      <link>https://halobench.com/log/#e-2026-08-14-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-2</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>On gfx1151, stock ROCm 3653e6d's own BF16 KV decode is SLOWER than its own F16 KV decode: 29.26 t/s vs 42.35 t/s at 32,768 depth (Qwen3.6-35B-A3B, fa on) — a 31% penalty for switching to the KV type the community recommends for its quality, with prefill roughly unchanged (587.4 vs 595.0). stew675's rdna-boosts branch (commit ed89854) fixes it with a native BF16 flash-attention tile kernel: 47.31 t/s BF16 decode, +61.7% over stock's own BF16 and stock's best config on this chip. — records: clm-0052</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] A benchmarking-harness defect, not a hardware limitation, produced this lab's earlier &quot;Nemotron-3-Super cannot allocate on ROCm at any depth&quot; verdict (clm-0050's first amendment, corrected 2026-08-14).</title>
      <link>https://halobench.com/log/#e-2026-08-14-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-3</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>A benchmarking-harness defect, not a hardware limitation, produced this lab's earlier &quot;Nemotron-3-Super cannot allocate on ROCm at any depth&quot; verdict (clm-0050's first amendment, corrected 2026-08-14). — records: clm-0053</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] 4 candidates → screened: npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-whisper</title>
      <link>https://halobench.com/log/#e-2026-08-14-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-4</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>4 candidates → screened: npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-whisper</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] qwen38-27b: listed → listed — RELEASED and CONFIRMED runnable.</title>
      <link>https://halobench.com/log/#e-2026-08-14-5</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-5</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>qwen38-27b: listed → listed — RELEASED and CONFIRMED runnable.</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 8 runs landed on cfg-0050, cfg-0051 (llama-bench)</title>
      <link>https://halobench.com/log/#e-2026-08-14-6</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-6</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>8 runs landed on cfg-0050, cfg-0051 (llama-bench) — records: cfg-0050, cfg-0051</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 44 runs landed on cfg-0052, cfg-0053, cfg-0056, cfg-0059, cfg-0060, cfg-0032, cfg-0062, cfg-0063, cfg-0064, cfg-0065 (llama-bench)</title>
      <link>https://halobench.com/log/#e-2026-08-14-7</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-7</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>44 runs landed on cfg-0052, cfg-0053, cfg-0056, cfg-0059, cfg-0060, cfg-0032, cfg-0062, cfg-0063, cfg-0064, cfg-0065 (llama-bench) — records: cfg-0052, cfg-0053, cfg-0056, cfg-0059, cfg-0060, cfg-0032, cfg-0062, cfg-0063, cfg-0064, cfg-0065</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 8 runs landed on cfg-0054, cfg-0055 (llama-bench)</title>
      <link>https://halobench.com/log/#e-2026-08-14-8</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-8</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>8 runs landed on cfg-0054, cfg-0055 (llama-bench) — records: cfg-0054, cfg-0055</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 2 runs landed on cfg-0057, cfg-0058 (perf-matrix-gtt120)</title>
      <link>https://halobench.com/log/#e-2026-08-14-9</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-9</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>2 runs landed on cfg-0057, cfg-0058 (perf-matrix-gtt120) — records: cfg-0057, cfg-0058</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 1 run landed on cfg-0061 (memgate-test)</title>
      <link>https://halobench.com/log/#e-2026-08-14-10</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-10</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>1 run landed on cfg-0061 (memgate-test) — records: run-0242</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 2 runs landed on cfg-0068, cfg-0069 (mtp-probe)</title>
      <link>https://halobench.com/log/#e-2026-08-14-11</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-14-11</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
      <description>2 runs landed on cfg-0068, cfg-0069 (mtp-probe) — records: cfg-0068, cfg-0069</description>
      <category>runs</category>
    </item>
    <item>
      <title>[claim] Quantised KV on gfx1151 splits three ways by build.</title>
      <link>https://halobench.com/log/#e-2026-08-13-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-13-1</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>Quantised KV on gfx1151 splits three ways by build. — records: clm-0051</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] 7 candidates → screened: gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, nemotron35-lightning-30b, qwen36-27b-mtp</title>
      <link>https://halobench.com/log/#e-2026-08-13-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-13-2</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>7 candidates → screened: gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, nemotron35-lightning-30b, qwen36-27b-mtp</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] nemotron35-lightning-30b: listed → acquired — Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b).</title>
      <link>https://halobench.com/log/#e-2026-08-13-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-13-3</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>nemotron35-lightning-30b: listed → acquired — Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b).</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 44 runs landed on cfg-0038, cfg-0039, cfg-0040, cfg-0041, cfg-0042, cfg-0043, cfg-0044, cfg-0045, cfg-0046, cfg-0047, cfg-0048, cfg-0049 (llama-bench)</title>
      <link>https://halobench.com/log/#e-2026-08-13-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-13-4</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>44 runs landed on cfg-0038, cfg-0039, cfg-0040, cfg-0041, cfg-0042, cfg-0043, cfg-0044, cfg-0045, cfg-0046, cfg-0047, cfg-0048, cfg-0049 (llama-bench) — records: cfg-0038, cfg-0039, cfg-0040, cfg-0041, cfg-0042, cfg-0043, cfg-0044, cfg-0045, cfg-0046, cfg-0047, cfg-0048, cfg-0049</description>
      <category>runs</category>
    </item>
    <item>
      <title>[gate] deepseek-v4-flash: listed → listed — Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context.</title>
      <link>https://halobench.com/log/#e-2026-08-12-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-12-1</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>deepseek-v4-flash: listed → listed — Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context. — records: clm-0031</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] nemotron35-lightning-30b: listed — Surfaced via an operator-shared r/AIDeveloperNews link (&quot;NVIDIA has launched Nemotron 3.5 Lightning&quot;) plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers.</title>
      <link>https://halobench.com/log/#e-2026-08-12-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-12-2</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>nemotron35-lightning-30b: listed — Surfaced via an operator-shared r/AIDeveloperNews link (&quot;NVIDIA has launched Nemotron 3.5 Lightning&quot;) plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers.</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 32 runs landed on cfg-0032, cfg-0033, cfg-0034, cfg-0035 (llama-bench)</title>
      <link>https://halobench.com/log/#e-2026-08-12-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-12-3</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>32 runs landed on cfg-0032, cfg-0033, cfg-0034, cfg-0035 (llama-bench) — records: cfg-0032, cfg-0033, cfg-0034, cfg-0035</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 16 runs landed on cfg-0036, cfg-0037 (llama-bench)</title>
      <link>https://halobench.com/log/#e-2026-08-12-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-12-4</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>16 runs landed on cfg-0036, cfg-0037 (llama-bench) — records: cfg-0036, cfg-0037</description>
      <category>runs</category>
    </item>
    <item>
      <title>[claim] A second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches.</title>
      <link>https://halobench.com/log/#e-2026-08-11-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-1</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>A second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches. — records: clm-0044</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038).</title>
      <link>https://halobench.com/log/#e-2026-08-11-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-2</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>The KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038). — records: clm-0045</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV.</title>
      <link>https://halobench.com/log/#e-2026-08-11-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-3</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>The community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV. — records: clm-0046</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Nemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567).</title>
      <link>https://halobench.com/log/#e-2026-08-11-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-4</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>Nemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). — records: clm-0047</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Simulator identity changes what tau2 measures.</title>
      <link>https://halobench.com/log/#e-2026-08-11-5</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-5</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>Simulator identity changes what tau2 measures. — records: clm-0048</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status.</title>
      <link>https://halobench.com/log/#e-2026-08-11-6</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-6</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>The 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status. — records: clm-0049</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] muse-glimmer-30b: blocked → screened — Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability).</title>
      <link>https://halobench.com/log/#e-2026-08-11-7</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-7</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>muse-glimmer-30b: blocked → screened — Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability).</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 9 runs landed on cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027 (tau2-bench-airline)</title>
      <link>https://halobench.com/log/#e-2026-08-11-8</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-8</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>9 runs landed on cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027 (tau2-bench-airline) — records: cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027</description>
      <category>runs</category>
    </item>
    <item>
      <title>[claim] The 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain.</title>
      <link>https://halobench.com/log/#e-2026-08-10-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-1</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>The 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain. — records: clm-0037</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Paired on identical tasks, q8_0 KV costs TURN EFFICIENCY: 228 turns against f16's 164 over the same 9 tasks, +39%, taking more turns on 6 of 9 and fewer on 1.</title>
      <link>https://halobench.com/log/#e-2026-08-10-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-2</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>Paired on identical tasks, q8_0 KV costs TURN EFFICIENCY: 228 turns against f16's 164 over the same 9 tasks, +39%, taking more turns on 6 of 9 and fewer on 1. — records: clm-0038</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] On reward the 122B and Nemotron are indistinguishable, but reward is the wrong headline: on tasks both get RIGHT, the 122B needs 19 turns and 2.0 minutes against Nemotron's 26 and 3.7 — 37% fewer loops and 85% less wall-clock to the same correct answer.</title>
      <link>https://halobench.com/log/#e-2026-08-10-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-3</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>On reward the 122B and Nemotron are indistinguishable, but reward is the wrong headline: on tasks both get RIGHT, the 122B needs 19 turns and 2.0 minutes against Nemotron's 26 and 3.7 — 37% fewer loops and 85% less wall-clock to the same correct answer. — records: clm-0039</description>
      <category>claim</category>
    </item>
    <item>
      <title>[superseded] by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models.</title>
      <link>https://halobench.com/log/#e-2026-08-10-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-4</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. — records: clm-0040, clm-0042</description>
      <category>superseded</category>
    </item>
    <item>
      <title>[claim] The KV dequant patch cuts energy 42% at 200k context — 146.1 Wh unpatched against 85.3 Wh patched for the same throughput benchmark — and patched q8_0 (85.3 Wh) beats f16 (89.6 Wh).</title>
      <link>https://halobench.com/log/#e-2026-08-10-5</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-5</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>The KV dequant patch cuts energy 42% at 200k context — 146.1 Wh unpatched against 85.3 Wh patched for the same throughput benchmark — and patched q8_0 (85.3 Wh) beats f16 (89.6 Wh). — records: clm-0041</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Per-task energy windows, matched to a common task set across arms, put a correct τ² answer at 6.81 Wh on the 122B with f16 KV, 9.48 Wh with q8_0, and 12.75 Wh on Nemotron.</title>
      <link>https://halobench.com/log/#e-2026-08-10-6</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-6</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>Per-task energy windows, matched to a common task set across arms, put a correct τ² answer at 6.81 Wh on the 122B with f16 KV, 9.48 Wh with q8_0, and 12.75 Wh on Nemotron. — records: clm-0042</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Every τ² arm was run with --user-llm set to the same model as --agent-llm, so cross-model comparisons changed the agent AND the user simulator together — exactly what the runbook forbids (&quot;hold both --user-llm and the judge fixed across comparisons, or results re-baseline silently&quot;).</title>
      <link>https://halobench.com/log/#e-2026-08-10-7</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-7</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>Every τ² arm was run with --user-llm set to the same model as --agent-llm, so cross-model comparisons changed the agent AND the user simulator together — exactly what the runbook forbids (&quot;hold both --user-llm and the judge fixed across comparisons, or results re-baseline silently&quot;). — records: clm-0043</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] 7 candidates → acquired: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b, muse-glimmer-30b, qwen36-27b-mtp</title>
      <link>https://halobench.com/log/#e-2026-08-10-8</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-8</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>7 candidates → acquired: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b, muse-glimmer-30b, qwen36-27b-mtp</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] 5 candidates → screened: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b</title>
      <link>https://halobench.com/log/#e-2026-08-10-9</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-9</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>5 candidates → screened: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] 16 candidates → listed: deepseek-v4-flash, gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, muse-glimmer-30b, npu-embeddinggemma, npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-qwen3-4b-thinking, npu-whisper, npu-whisper, qwen38-27b, qwen38-27b</title>
      <link>https://halobench.com/log/#e-2026-08-10-10</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-10</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>16 candidates → listed: deepseek-v4-flash, gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, muse-glimmer-30b, npu-embeddinggemma, npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-qwen3-4b-thinking, npu-whisper, npu-whisper, qwen38-27b, qwen38-27b — records: clm-0031, clm-0034, clm-0015, clm-0018</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] 4 candidates → blocked: laguna-s-21, laguna-s-21, muse-glimmer-30b, muse-glimmer-30b</title>
      <link>https://halobench.com/log/#e-2026-08-10-11</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-11</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>4 candidates → blocked: laguna-s-21, laguna-s-21, muse-glimmer-30b, muse-glimmer-30b</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 2 runs landed on cfg-0026, cfg-0027 (tau2-bench-airline)</title>
      <link>https://halobench.com/log/#e-2026-08-10-12</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-12</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>2 runs landed on cfg-0026, cfg-0027 (tau2-bench-airline) — records: cfg-0026, cfg-0027</description>
      <category>runs</category>
    </item>
    <item>
      <title>[superseded] the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors.</title>
      <link>https://halobench.com/log/#e-2026-08-09-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-1</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. — records: clm-0030, clm-0037</description>
      <category>superseded</category>
    </item>
    <item>
      <title>[claim] A community DeepSeek-V4-Flash report independently confirms our mmap/GTT double-residency finding, demonstrates a 120 GiB GTT ceiling in production use, and — most consequentially — reports Vulkan BEATING ROCm on 3 of 4 cells including 55% faster decode at depth.</title>
      <link>https://halobench.com/log/#e-2026-08-09-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-2</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>A community DeepSeek-V4-Flash report independently confirms our mmap/GTT double-residency finding, demonstrates a 120 GiB GTT ceiling in production use, and — most consequentially — reports Vulkan BEATING ROCm on 3 of 4 cells including 55% faster decode at depth. — records: clm-0031</description>
      <category>claim</category>
    </item>
  </channel>
</rss>
