Home › Evidence › Records › clm-0057

clm-0057

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0057/ — this address never moves; the anchor /records/#clm-0057 keeps resolving

Qwen3.8-27B met its registered agentic prediction: on the standing 5-task tau2 airline smoke subset it scored 1.000 (5/5, 21 tool-call messages, 0 empty assistant turns, VALID under the smoke gate) against the >= 0.80 bar recorded in the candidate record on 2026-08-14, BEFORE any measurement — and against the incumbent qwen36-27b-mtp's 0.80 on the identical tasks under the identical pinned-simulator protocol, where the incumbent failed task 2 and this model passed it. At n=5 one task is worth 0.20: that is a margin met, not a significance claim. The full run behind it: mean 0.577 over an effective n=26 of the 50-task set (the run hit its 11,400 s wall bound at rc=124; all 26 completed tasks scored, 165 tool-call messages, 0 empty turns, 0 max-steps cuts), on ROCm Q8_0 with draft-mtp n_max=3, reasoning_effort pinned to medium (positive-control verified on the baked template), agent max_tokens 4096, pinned haiku-4.5 simulator, seed 42. The first five tasks are the easy end of the set — this model's own full-run mean is 0.577 against 1.000 on them — which is exactly why smoke and full are reported as separate numbers rather than blended. The max_tokens 4096 output cap is a fingerprint deviation: the historical tau2 series ran uncapped, so 0.577 is not directly comparable to those means.

verified 2026-08-15 · volatility medium
evidence run-0267 run-0271

Note — the record's own working

METHOD — tau2-bench airline, harness 668d3bc, --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0, Anthropic-only routing). Agent: local llama-server endpoint on cfg-0066 (ROCm Q8_0, draft-mtp n_max=3 under the clm-0055 hard cap, -c 32768, --parallel 1, --load-mode none), agent sampling temperature 1.0 / top_p 0.95 / max_tokens 4096, and reasoning_effort pinned to medium via --chat-template-kwargs — llama.cpp discards the OpenAI reasoning_effort request field (PR 26941 pending), so the template kwarg is the only route that reaches every request, and the queue PROVED it lands before any task ran (xhigh and medium renders differ; the medium render carries no xhigh; run-0267's comment). Q8_0 is outside protocol.json quants.expected; the override is recorded in run-meta per the protocol's override discipline (Q8_0 is the candidate record's designated primary screening quant, chosen to remove the quant confound from a capability screen). THE PREDICTION — recorded on the candidate yaml 2026-08-14, before release-day measurement: "Qwen3.8-27B meets or beats 0.80 on the same smoke" (the vendor's headline claims are all agentic, which is why the smoke was pre-registered as the independent check). Outcome: MET, 1.000 vs the bar of 0.80. The incumbent comparison is same-tasks, same-protocol, same-simulator (qwen36-27b-mtp re-screen of 2026-08-13, recorded in that candidate's history): 0.80, failing task 2 — the task this model completed in 576 s, its longest smoke task. THE FULL RUN — 26 of 50 tasks inside the hard 11,400 s bound (tasks arrive in id order, so the effective set is tasks 0-25, not a random draw; per-task rewards, turns and wall times are in run-0271's trace). Passed 15/26. Median turns 24, p90 36, max 44, none cut at max_steps. Wall time per task ranged 50 s to 1,176 s. SMOKE VALIDITY GATE (protocol, added 2026-08-14): tool_call_messages recorded alongside both means — 21 (first five) and 165 (all 26), zero empty assistant turns, zero infrastructure errors, so neither number is a do-nothing artifact. CAVEATS — (1) the output cap: max_tokens 4096 bounds each assistant turn; historical uncapped arms allowed longer reasoning, so cross-series comparison of the 0.577 must carry this asterisk (the smoke-vs-incumbent comparison is NOT affected the same way — the incumbent smoke ran under its own screen-tier settings, and both are n=5 cliff-detectors, not rankings). (2) Effective-n truncation: 26 of 50 by wall bound is a completed-prefix comparison in clm-0037's sense if set against a full-50 mean; compare only matched prefixes. (3) Speed is the cost of the capability: ~18 t/s decode under n_max=3 speculation (clm-0055) put the median passing task at several minutes of wall time — the capability verdict and the latency verdict are separate findings and the candidate record carries both.

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b