HomeMethod › Corrections

Corrections

Every claim this lab has withdrawn, kept at its citable id with its full original text — struck through nowhere, hidden never. Each entry names what replaced it; the worked diagnosis of what went wrong lives on the record itself. These entries also appear inline in the Log, badged, in the same timeline as every other learning: this page is a generated filter of that stream, not a shrine.

4 corrections on record · a headline result that gets retracted is the process working — catching it before it became policy is the retraction's whole value

clm-0050correctedmeasured-heremed ●●○verified 2026-08-14

AMENDED 2026-08-14 — the rule below is PER-MODEL AND PER-PHASE, not fleet-wide; see the correction history and the amendment note. On gfx1151 at f16 KV, stock Vulkan beat stock ROCm in EVERY cell of a matched matrix measured on Qwen3.6-35B-A3B (one binary commit 3653e6d, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k. That finding was correct as measured and remains the reference case for this model. It does NOT generalise: the 2026-08-14 perf-matrix sweep found gpt-oss-120b's PREFILL inverts hard in ROCm's favour at depth — Vulkan −52% at d65536, −84% at d131072 (23.1 t/s vs ROCm's 144.3, tight across 3 reps) — while gpt-oss-120b's own DECODE still favours Vulkan (+10 to +16%), same direction as Qwen3.6-35B. Nemotron-3-Super is now a THIRD supporting data point rather than an absence: the OOM first recorded against it (twice, including at an expanded 122 GiB GTT ceiling) was a benchmarking-harness artifact of running stock ROCm under mmap on a model within ~40 GiB of the box's full memory, corrected 2026-08-14 — see clm-0053. On the identical binary with `--load-mode none`, Nemotron-3-Super UD-Q4_K_M at d32768 measures 244.05 pp / 16.43 tg (median of 3 post-reboot reps; an earlier single-rep check the same day landed within 2% at 248.79 pp / 16.39 tg) against Vulkan's own 139.8 pp / 17.56 tg at the same cell: ROCm +74.6% prefill, Vulkan +6.4% decode — the same prefill-favours-ROCm, decode-favours-Vulkan split gpt-oss-120b shows, now measured on a third architecture. The r/LocalLLaMA claim that ROCm leads Vulkan 3.5x at 65k depth — measured on an RDNA2 V620 — still inverts on gfx1151 for Qwen3.6-35B decode/prefill and for gpt-oss-120b decode, but NOT for gpt-oss-120b's own prefill, which is the live counter-example. The community fork (v0.6.1) over stock Vulkan remains a prefill-only win at f16 KV on Qwen3.6-35B: +13/+13/+1/+2.5/+16% by depth, decode unchanged (±1%), and no stride bug (41/41 layers on GPU, CPU utilisation identical to stock) — untested on the other two models.

corrected by: clm-0053 — the withdrawn record keeps its URL and full text; the successor carries the number to cite

  • 2026-08-14

    Original text ("stock Vulkan beats stock ROCm in EVERY cell") was true of its one-model matrix but read as a fleet-wide backend rule. The 2026-08-14 perf-matrix sweep (gpt-oss-120b, Nemotron-3-Super) falsified the fleet-wide reading: gpt-oss-120b prefill favours ROCm by a wide and growing margin at depth (opposite direction from Qwen3.6-35B and from gpt-oss-120b's own decode). Rule restated as per-model, per-phase. The original Qwen3.6-35B numbers are unchanged and correct; only the scope of the claim was wrong.

  • 2026-08-14

    The Nemotron-3-Super OOM behind this claim's first amendment ("favours Vulkan only in the degenerate sense") was itself a benchmarking-harness artifact, not a hardware limitation: the perf-matrix and perf-matrix-gtt120 llama-bench queues that produced run-0234/run-0235 never carried --load-mode none, so stock ROCm was benchmarked under mmap on a model within ~40 GiB of the box's full 122 GiB — full diagnosis in clm-0053. Re-run with --load-mode none on the identical binary, Nemotron-3-Super runs cleanly on ROCm (run-0236, run-0237: 244.05 pp / 16.43 tg, median of 3 post-reboot reps at d32768; an earlier single-rep check the same day, run-0243/run-0244, landed within 2% at 248.79 pp / 16.39 tg) and becomes a THIRD supporting data point for the per-model/per-phase rule rather than an absence: ROCm +74.6% prefill / Vulkan +6.4% decode against Vulkan's own run-0190/run-0191 (139.8 pp / 17.56 tg) — the same prefill-favours-ROCm, decode-favours-Vulkan split gpt-oss-120b shows. The per-model, per-phase rule itself is unchanged; only Nemotron's status in it moves from absent to confirming.

clm-0040supersededmeasured-herelow ●○○verified 2026-08-10

SUPERSEDED by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. As measured here: a correct τ² answer cost 78.0 Wh on the 122B with f16 KV, 100.1 Wh on Nemotron, and 115.3 Wh on the 122B with q8_0 KV, but only 8-27% of each arm's wall time fell inside a scored task. aihydra's power envelope stands on its own: idle 10.1 W, 150-168 W under inference, peaking at 218 W.

corrected by: clm-0042 — the withdrawn record keeps its URL and full text; the successor carries the number to cite

clm-0030supersededmeasured-herelow ●○○verified 2026-08-09

SUPERSEDED: the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. The sustained score across a realistic sample is clm-0037's 0.545 (n=22), which is the number to cite for this model on tau2 airline. This run was nonetheless the project's first genuine capability measurement rather than a throughput number, and it proved the harness and scoring path work.

corrected by: clm-0037 — the withdrawn record keeps its URL and full text; the successor carries the number to cite

clm-0035retractedmeasured-herelow ●○○verified 2026-08-09

RETRACTED: the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. The corrected 122B score is clm-0037's; the corrected cross-model comparison is clm-0039's. What survives is categorical, not scored: the 122B fails to terminate some tasks with thinking on, and gpt-oss fails the domain outright.

corrected by: clm-0037 clm-0039 — the withdrawn record keeps its URL and full text; the successor carries the number to cite