clm-0111
On Qwen3.5-122B-A10B UD-Q4_K_M (build ce7689f, llama.cpp-kvfix tree, ROCm0 Radeon 8060S gfx1151 122880 MiB, q8_0/q8_0 KV baseline) the production-optimisation matrix interior conclusions hold (internal comparison, same model/quant/build): (1) interactive first-token latency (ttft) is flat at ~1420 ms across batch 512-2048 at ub=512 (all N=3, CV<1.6%), and degrades when ubatch drops below 512 — ub256 ttft ~1683-1691 ms, ub128 ttft ~2364 ms — so ub=512 is the interactive optimum and tpot stays flat ~45.8 ms/tok throughout. (2) Depth d262144 LOADS with q8_0 KV, no OOM; pp16384 = 62.25 tok/s, tg256 = 8.448 t/s — the >=200k production-depth probe is satisfied on this config. (3) KV q4_0 is NOT statistically equivalent to f16: ttft +0.14% (t=-0.39, within noise) but decode (tpot) +0.407% is resolvable (t=-21.45, N=3 CV<0.04%) — q4_0 hides a small but measurable decode regression and is UNPROFITABLE at this depth; it is near-free as the protocol expects on a hybrid-attention minority but not free. Boundary: internal matrix only — same model, same quant UD-Q4_K_M, same build ce7689f, same backend ROCm0; only batch/ubatch, KV quant, and context depth vary. NO cross-model, cross-build, or cross-backend-equivalence claim is made or inherited.