llm-lab bench campaign · 2x RTX PRO 4000 Blackwell

Qwen3.8-27B vs Qwen3.6-35B

Campaign closed Exhaustive head-to-head complete. Verdict based on matched evidence.

Last updated 2026-08-16 · SGLang Phase 1 final: NVFP4 2-GPU decode 57.32 t/s without speculation and 96.22 t/s with MTP (+68%), 1-GPU did not fit; the MTP arm is about 3x the llama.cpp Q4 serve but remains about 29% below production near 135 t/s, so Phase 2 stays off · addendum: Opus 4.6 now has quick-quality numbers via the same first-party claude -p path - corrected composite 94.4, the best on file (flags carried), knowledge 86.0 tops the table, needle 100 after a harness-cap artifact was diagnosed and remeasured · full Opus per-task tier detail remains its own section · CAMPAIGN CLOSED · Final verdict: keep production Qwen3.6-35B; Qwen3.8-27B lab-only (fails all three promotion clauses) yet the strongest quality candidate ever benched - matched agentic cell 42 vs 41 at the measured suite ceiling, best expert tier 16/18, tool-calling 90.6, maths +5.2, MTP +58% · Q8_0 quant arm final: Q4_K_M +56% decode, ~10.5 GB lighter · production restored and verified live (health 200) · eighteen sources on file

Production, defending

Qwen3.6-35B-A3B

Mixture-of-experts, 35B total / 3B active · Q4_K_M · the incumbent serving all live workloads

Candidate, challenging

Qwen3.8-27B

Dense 27B, thinking-native, coding + agentic positioning · Q4_K_M · released 2026-08-15

Quant-matched at Q4_K_M on identical hardware, both single-GPU and dual-GPU (layer-split) configurations. Same frozen eval sets, same serving stack.

Scoreboard

One cell per axis, each with its final campaign read. Filled tag = proven separation. Outlined tag = measured leader whose gap remains inside the noise band (same-cell swing up to ~2.6 points on quality evals); the promotion verdict does not depend on those noise-band leads.

Quick quality
91.7 / 93.8 vs 92.1 / 92.6 · matched budget
Qwen3.8 +0.4
Speed · decode
~32 Q4 / 96.22 SGLang MTP vs ~135 t/s
Qwen3.6 ~1.4x vs best arm
Refusal behavior
0.00/0.07 vs 0.14/0.21 refusal rate
Qwen3.8 wins
Card-bench
VRAM to 3.8 · speed + power to 3.6
Qwen3.6 on net
Tool calling /100
90.6 vs 84.4
Qwen3.8 +6.2
Agentic suite
think-off 42 vs 41 /45 · 71.1 vs 83.5 /100
Qwen3.8 matched cell
Token efficiency
766/906 vs 674/725 tok/item
Qwen3.6 leaner
Thinking-mode arms
35 / 39 / 39 / 42 across the ladder
Off wins · ladder complete
Frontier calibration · Opus 4.6
expert 14/18 · hard 12/12 · ≈41/45 - production's exact tier profile
Suite ceiling ≈ 41-42/45

Head to head

Every measured metric, one row each. Cells show both GPU configurations as 1 GPU / 2 GPU. Every row names its winner: a filled tag is a proven separation, an outlined tag is the current leader while the gap is still inside the noise band (winner judged on the mean of both configs, ~2.6-point same-cell swing on quality evals). Dead heat means the scores are literally identical. A purple-tinted Opus cell marks the two rows where Opus 4.6 beats both Qwens outright (composite and knowledge); on every other row at least one local model matches or beats the frontier reference.

Metric (1 GPU / 2 GPU)Qwen3.8-27BQwen3.6-35BReadOpus 4.6 (frontier ref)
Quick quality · frozen suite (matched 16K reasoning budget)
Composite91.7 / 93.892.1 / 92.6Qwen3.8 +0.4 local · Opus best on file94.4 beats both Qwenscorrected · the run as printed said 69.4 because the needle category scored 0 on a harness-cap artifact (diagnosed and remeasured, see the needle row); corrected composite = (86.0 + 96.7 + 95.0 + 100.0) / 4 · best composite on file, flags below apply
Knowledge80.0 / 80.080.0 / 82.0Qwen3.6 +1.0 local · Opus 86 tops the table86.0 beats both Qwensbest on file · contamination-exposed for any frontier model
Maths97.0 / 100.093.3 / 93.3Qwen3.8 +5.2 · edges Opus by 1.8 on the mean96.70 errors · tools disabled, model-only arithmetic
Instruction-following90.0 / 95.095.0 / 95.0Qwen3.6 +2.5 local · Opus ties production at 9595.0
Long-context needle100.0 / 100.0100.0 / 100.0Three-way 100100.0remeasured · the first pass scored 0/6 at the suite's 64-token output cap because Claude Code treats its output cap as a hard ERROR, not a truncation point - and of the two answers that did return, one was harness identity confusion and one was Opus finding the buried needle and refusing to relay it as a suspected prompt injection; with the cap floored at 512 all six extractions came back exact
Refusal behavior · 14 calibrated probes per cell
Refusal rate0.00 / 0.070.14 / 0.21Qwen3.8 wins local · Opus also 0.000.00
Answered accuracy1.00 / 0.900.89 / 1.00Wash local · Opus 0.36 flagged0.3611 of 14 probes are accuracy-scored · Opus answered everything but agreed with the probe answer key far less often than the locals - with no temperature control and a different system prompt this is flagged, not concluded
Speed probe · 2K context, 3-run median after warmup
Decode (t/s)31.4 / 32.8139.7 / 142.0Qwen3.6 4.4xn/a · cloud
Prefill (t/s)844 / 8731057 / 1048Qwen3.6 1.2xn/a · cloud
TTFT p50 (ms)611 / 591488 / 493Qwen3.6 winsn/a · cloud
SGLang Phase 1 · NVFP4 weights, bench port, 2-GPU tensor parallel · Phase 2 off
Decode, no-spec / MTP (t/s)1 GPU: no fit · 2 GPU: 57.32 / 96.22MTP +68% over the SGLang no-spec arm and about 3x the candidate's llama.cpp Q4 serve~135production Qwen3.6-35B MoEQwen3.6 ~1.4x vs MTPn/a · cloud
Prefill, no-spec (t/s)1 GPU: no fit · 2 GPU: 22,283.5not matched in this Phase 1 probeSGLang wins candidate prefilln/a · cloud
Card-bench · 50%-of-context prompts, 3-run median · both configs in
Decode at 8K context (t/s)30.6 / 31.9128.4 / 130.1Qwen3.6 4.1xn/a · cloud
Decode at 131K context (t/s)21.2 / 22.6n/a / 77.0Qwen3.6 3.4xn/a · cloud
Concurrent 4-stream aggregate (t/s)68.9 / 77.1243.8 / 252.4Qwen3.6 3.4xn/a · cloud
VRAM peak at 8K (MiB)16,834 / 17,61621,504 / 22,034Qwen3.8 winsn/a · cloud
Power draw at 8K (W)150.4 / 224.8127.8 / 165.6Qwen3.6 winsn/a · cloud
Energy per token at 8K (J/tok)4.9 / 7.11.0 / 1.3Qwen3.6 ~5xn/a · cloud
KV-cache precision A/B · candidate weights, same 3-run methodology · lab default is q8_0 KV (a this-campaign discovery: every published card-bench number is a q8-KV number)
Decode at 131K by KV type (t/s, 1 / 2 GPU)f16: crash / 24.9 · q8: 21.2 / 22.6 · q4: 20.9 / 22.3n/a · diagnosticf16 +10% where it fitsn/a
Prefill at 131K by KV type (t/s, 2 GPU)f16: 49.7 · q8: 41.2 · q4: 42.3n/a · diagnosticf16 +21%n/a
VRAM peak at 131K by KV type (MiB, 2 GPU)f16: 26,496 · q8: 23,596 · q4: 21,548n/a · diagnosticq4 wins, -4.8 GB vs f16n/a
131K on a single 24 GB cardf16 KV does not fit (server crash) · q8: 21.5 GiB ok · q4: 19.0 GiB okn/a · diagnosticquantised KV requiredn/a
Weight-quant A/B · Q4_K_M vs Q8_0 on the same weights · 2-GPU only (Q8_0 is 27.1 GiB and cannot fit one card), same 3-run methodology
Decode at 8K by weight quant (t/s, 2 GPU)Q4_K_M: 31.9 · Q8_0: 20.4n/a · diagnosticQ4 +56%n/a
Decode at 131K by weight quant (t/s, 2 GPU)Q4_K_M: 22.6 · Q8_0: 16.2n/a · diagnosticQ4 +40%n/a
Prefill at 131K by weight quant (t/s, 2 GPU)Q4_K_M: 41.2 · Q8_0: 37.7n/a · diagnosticQ4 +9%n/a
VRAM peak by weight quant (MiB, 2 GPU, 8K / 131K)Q4_K_M: 17,616 / 23,596 · Q8_0: 28,396 / 34,376n/a · diagnosticQ4 saves ~10.5 GBn/a
Tool calling · 51 deterministic scenarios, temperature 0
Tool-calling /10090.684.4Qwen3.8 +6.2not run
Agentic · complete (64K context, matched to production's reference)
Agentic suite /100 (weighted)71.183.5Qwen3.6 +12.4n/a · harness-specific
Agentic 64K context /45 (shared 4096-token output budget)3541Qwen3.6 winsn/a · no cap applies
Agentic /45 at raised 16K output cap (diagnostic, reported separately)39expert 15/18 (was 11/18) · hard 9/12 (hard_base 0/3 by clean 630s timeouts) · medium/small/trivial 15/1542on file, same capQwen3.6 winsn/a · no cap applies
Agentic /45 with thinking OFF at the matrix 4096 cap (matched cell)42FINAL · expert 16/18 (thinking-on matrix: 11/18) · hard 11/12 · medium/small/trivial 15/15 · zero truncations, zero loops41measured 08-15 19:22 in the identical config (thinking off, 4096 cap, 64K ctx) · expert 14/18: expert_calc 1/3, expert_money 1/3 · hard 12/12 · medium/small/trivial 15/15 · production's best anywhere on file stays 42 (thinking on, 16K cap)Qwen3.8 wins the matched cell≈41same 19 tasks, hidden oracles, via first-party claude -p · full 3-run cells now on every expert and hard task: expert 14/18 (expert_calc 1/3, expert_money 1/3 - production's exact failure profile; the other four experts all 3/3), hard 12/12, all 24 tier-repeat runs passed at 12-34 s wall · with single-run medium/small/trivial 15/15 that lands ≈41/45 - identical to production's split, one under the candidate's 42
Agentic /45 with thinking capped at 2048 (diagnostic)39FINAL · expert 15/18 (expert_money 2/3 where thinking-off went 3/3) · hard 9/12: hard_base 0/3, all three 630s wall-clock timeouts mid-work, no loops - per-turn thinking starves the time budget on the longest-horizon task (thinking-off went 2/3) · medium/small/trivial 15/1542same comparators as the row aboveQwen3.6 winsn/a · no cap applies
Token efficiency · quick-quality suite, output tokens per scored item (matched 16K reasoning budget)
Tokens per item (all scored)765 / 906674 / 725Qwen3.6 ~16% leanern/a · not comparable
Tokens per correct answerearlier 633 vs 768 candidate lead withdrawn - it compared a capped run against an uncapped onesuperseded by the matched all-scored row aboveCap artifactn/a
Tokens-to-answer, agentic tasksnot run - the matched-cap quick-quality row above settled the token-efficiency axis (production ~16% leaner)not runClosed unrunn/a

The matched-budget quality repeats are complete: the candidate leads the local composite by +0.4 on the mean, inside the measured ~2.6-point noise band, while production is ~16% leaner per scored item. Card-bench gives the candidate a clean footprint win - 4.3-4.7 GiB less VRAM at 8K on both configs - but production keeps 3.4-4.1x speed and roughly 5x better energy per token, and every candidate cell ran throttled against the software power cap. On tool calling the candidate's +6.2 comes from a deterministic temperature-0 suite (45 of 51 scenarios passed), so it is tagged as a win, not noise; its six misses cluster in trap scenarios, including one that matters operationally - it fired a destructive delete with confirm=true without asking the user first. The speed gap is the expected dense-vs-MoE cost: minor for chat (32 t/s outruns reading speed), compounding for thinking chains and agentic loops. The Opus 4.6 column is a frontier reference, not a contender: it ran the same 19 tasks against the same hidden oracles via first-party claude -p on cloud hardware, so pass/fail transfers but speed, VRAM, power, and token metrics do not. Quick-quality has now been run for Opus through the same first-party path (a driver that swaps the suite's chat call for a claude -p subprocess), and its numbers carry their own flags: the Claude Code harness system prompt is injected, there is no temperature or seed control (locals ran temp 0, seed 26), tools are disabled, knowledge and maths are contamination-exposed for any frontier model, and the needle category's output cap had to be floored at 512 tokens against the locals' 64 - output-budget parity is impossible below 512 on this path. Its agentic result is the calibration that reframes the whole /45 axis: 17/19 on the first pass, failing only the exact two expert tasks production drops, with repeats pinning both at 1/3 - so the suite's effective ceiling is about 41-42/45 and both local finalists are already sitting on it.

Opus 4.6 reference: per-task detail on the discriminating tiers

The frontier reference, opened up. Every expert and hard task carries a full 3-run Opus cell (24/24 tier-repeat runs completed; the two failure cells were repeated in the earlier calibration pass). Comparators are the matched thinking-off cells for both local models. Medium, small, and trivial tiers were 15/15 for all three models and are collapsed into one row; Opus ran those once, not three times. Wall-clock is not comparable across the columns - Opus ran on cloud hardware through a first-party CLI harness (repeat runs landed in 12-34 seconds per task) - so this table reports pass counts only.

TaskQwen3.8-27B (think-off)Qwen3.6-35B prod (think-off)ReadOpus 4.6
Expert tier · 6 tasks × 3 runs
expert_calc1/31/3Nobody clears it1/3
expert_money3/31/3Only Qwen3.8 clears it1/3
expert_csv_dialect3/33/3Three-way clean3/3
expert_debug_eventbus3/33/3Three-way clean3/3
expert_reservations3/33/3Three-way clean3/3
expert_scheduler3/33/3Three-way clean3/3
Expert tier /181614Best expert tier on file14
Hard tier · 4 tasks × 3 runs
hard_base2/33/3Qwen3.8 gives one back3/3
hard_duration3/33/3Three-way clean3/3
hard_intervals3/33/3Three-way clean3/3
hard_lru3/33/3Three-way clean3/3
Hard tier /121112Opus matches production12
Lower tiers · medium + small + trivial
Medium / small / trivial /151515Saturated for all three15single-run for Opus
Total /454241Qwen3.8 by one, inside noise≈41

The reportable headline: a frontier model, run over the same tasks, sandbox, and hidden oracles, produces production's tier profile exactly - down to failing the same two expert tasks at the same 1/3 rate - and lands one point under the local candidate's matched cell. The two structural cells are expert_calc, which no model clears at better than 1/3 in any configuration (task-hard, effectively excluded from the reachable ceiling), and expert_money, where Qwen3.8 thinking-off is the only model of the three to score above 1/3. That is why the suite's effective ceiling is 41-42/45 and why the /45 axis stops discriminating at the top. What this table cannot say: anything about speed, cost, or tokens (different harness and hardware). Quick-quality has since been run for Opus over the same first-party path - those numbers, with their flags, live in the head-to-head table above.

Cross-check against independent testing

Eighteen external sources so far. Source 1: a published 10-task head-to-head of Qwen 3.8 27B vs Qwen 3.6 27B (dual RTX 6000 Pro, BF16, multi-token prediction enabled). Source 2: a live 34-entry community leaderboard running its own task suite across 3090 and DGX Spark clusters. Source 3: a context-depth decode sweep of Qwen 3.8 27B NVFP4 on 4x RTX 3090, comparing FP8 vs BF16 KV cache out to 254K. Source 4: llama-bench decode numbers for the same GGUF quants on Turing and Ada workstation cards - the closest stack to ours. Source 5: a ts-bench coding-agent run of the UD_Q4_K_XL GGUF with thinking off, on llama.cpp like ours, plus the same author's follow-up forensic trace of its loop failures. Sources 6 and 7: serving recipes from one author (DGX Spark + RTX 6000 PRO on vLLM, then a single RTX 5090 on SGLang), configuration detail but no benchmarks - the 5090 repo has since added a work-in-progress banner and a worked 32GB memory budget, still without a single measured number. Source 8: a single-DGX-Spark deployment under a vLLM 0.27 nightly, and the first repo to publish hard multi-token-prediction numbers for this model. Source 9: the vendor's own published benchmark table (via a widely shared post) - marketing numbers at full precision, held to a lower evidentiary standard than the independent sources. Source 10: the official vLLM serving recipe for the model - authoritative configuration, no benchmarks. Source 11: a measured MTP speculative-decoding A/B on llama.cpp with the same unsloth Q4_K_M GGUF we run, on 24GB cards. Source 12: a published CRUDbench comparison (X, @LottoLabs, 2026-08-15) of the model's non-thinking vs highest-effort thinking modes on a tightly specified deletion task. Source 13: a community NVFP4 quantization with the MTP draft head preserved in bf16, with measured vLLM throughput across concurrency levels on four 16GB Blackwell cards. Source 14: relayed vendor deployment guidance on reasoning-effort levels and long-context extension. Source 15: an Intel Arc Pro B70 inference cookbook with a measured MTP depth ladder (vLLM XPU, GPTQ-Int4 weights, BF16 MTP head) - self-reported, held provisional. Source 16: a thinking-budget sweep on an HTML canvas generation test (X, 2026-08-16) - qualitative video comparison, no scores. Source 17: a practitioner report preferring a medium thinking budget standalone and no-thinking execution under an external planner (X, @TheAhmadOsman, 2026-08-16) - anecdote. Source 18: the official SGLang Qwen3.8-27B deployment cookbook, including the exact H200 FP8 balanced cell and architecture-specific guidance - authoritative configuration, no published speed table. Different hardware and precision: ratios and direction should transfer even where absolutes cannot.

Their findingOur measurementVerdict
Source 1 · 10-task head-to-head, dual RTX 6000 Pro, BF16 + MTP
3.8 wins 9/10 real tasks vs 3.6-27BDirection agrees at matched budget on quick quality (91.7/93.8 vs 92.1/92.6, candidate +0.4 mean) and the agentic arbiter is now in: thinking-off the candidate takes the matched cell 42 to 41Consistent
3.8 decodes 1.9x faster than 3.6-27B (93 vs 50 t/s)Plain llama.cpp Q4 is effectively tied with its 3.6-27B sibling on our stack, but the mechanism transfers: our llama.cpp MTP arm reaches 49.8 t/s (+58%), and SGLang NVFP4 MTP reaches 96.22 t/s on 2 GPUs. Against our faster 35B MoE production model near 135 t/s, the candidate still trails by about 29%Mechanism reproduced
3.8 uses 2.9x more tokens per taskDirection now reproduces at matched budget, at far smaller magnitude: 3.8 spends ~16% more tokens per scored item than production (766/906 vs 674/725). The earlier "3.8 leaner" read was an artifact of the candidate's tighter output cap. His 2.9x came from long agentic jobs, where the thinking channel compounds the appetiteReproduced in direction
Source 2 · 34-entry live leaderboard, 3090 / DGX Spark clusters, vLLM + llama.cpp
3.8-27B (INT4, 2x 3090) is their overall #1 on quality (97.1), above every Qwen3.6 variant, MiMo-V2.5, GLM-5.2 and DeepSeek-V4 FlashSame direction as our quality result, stronger magnitude. Their metric weights quality 70/30 over speed, so it structurally favors 3.8. Our completed agentic /100 comparison did not support promotionQuality direction consistent
Qwen3.6-35B-A3B still owns raw decode on their boards (up to 178 t/s)Matches our proven speed split: production 3.3-4.4x faster on every throughput axisConsistent
Their NVFP4 build of 3.8 scores 0.0 with 69/69 task failuresWe run GGUF Q4_K_M, unaffected. Flag: the NVFP4 quant path for this model looks broken right now, so no NVFP4 arm until that maturesNoted, avoided
Source 3 · 72-point context-depth sweep, 4x RTX 3090, NVFP4 weights, FP8 vs BF16 KV cache
FP8 KV cache holds decode nearly flat at depth: -2.4% at 254K vs -58% with BF16 KVA/B CLOSED on our stack, and the direction is the opposite of their claim. Card-bench's default was already q8_0 KV, so the 30.6-to-21.2 depth slide happens WITH quantised KV. The chained true-f16 baseline then measured FASTER at depth, not slower: f16 KV decodes 24.9 t/s at 131K on 2 GPUs vs q8's 22.6 (+10%) and prefills 49.7 vs 41.2 (+21%) - on our llama.cpp path, quantising KV costs depth speed, it does not recover it. What quantised KV buys is fit: f16 KV cannot hold 131K on one 24 GB card (server crash, 25.9 GiB peak on 2 GPUs), and q4_0 saves a further ~2 GB over q8 at no meaningful speed cost. Most of the depth slide is attention compute either way - even f16 falls 32.1 to 24.9 from 8K to 131KReversed on our stack
3.8-27B NVFP4 serves fine via a Marlin kernel, even on GPUs with no native FP4 or FP8Counter-evidence to Source 2's broken NVFP4 entry, now reinforced by our own clean SGLang Phase 1 serve. Both are speed evidence only, not quality evidence, so the NVFP4 checkpoint still cannot support a promotion claimServe path reproduced
Source 4 · llama-bench tg128, same GGUF quants, Turing 24GB (SM75) vs Ada 48GB (SM89)
Q4_K_M decodes 24.2 t/s on a 2018 Turing card and 46.0 on AdaOur 31-33 t/s sits exactly where our card class belongs on that ladder, on the same engine and quant. Confirms the 4x deficit vs production is the dense-27B architecture, not a serve misconfiguration on our sideConsistent
Dual-GPU layer split buys nothing: 46.5 vs 46.0 t/s single-cardMatches our split exactly (31.9 vs 30.6 across configs, within noise). Third stack to show single-stream decode does not scale with a second cardConsistent
Q8_0 (26.6 GiB) does not load on a 24GB card; their 256K recipe is Q6_K plus 4-bit KV cacheThe fit prediction transferred: our 27.1 GiB Q8_0 arm was 2-GPU only. It also settled the tradeoff in Q4_K_M's favour: Q4 decoded 40-56% faster and used about 10.5 GB less VRAM. Quantised KV remained the fit lever at 131K, where true f16 KV could not fit one 24 GB cardFit prediction reproduced
Source 5 · ts-bench coding-agent run, UD_Q4_K_XL GGUF, thinking off, llama.cpp
3.8-27B is significantly worse as a coding agent than 3.5 or 3.6, getting stuck in infinite loops until automatic compactionSame direction as our decided agentic axis (35/45 vs production's 41/45), and the same signature: our harness tagged every failed stall "looped" before the trace showed they are 4096-token cap truncations. A second llama.cpp stack seeing agent loops that 3.5/3.6 never produced makes the weakness look real, not localConsistent
Their loops happen with thinking OFFNot reproduced on our stack. The completed thinking-off arm scored 42/45 with zero truncations and zero loops, versus 35/45 with thinking on at the matched cap. Their long-context loop hazard remains useful external evidence, but it did not predict our best serving modeNot reproduced locally
They initially suspected a chat-template error or llama.cpp optimisation issue rather than the modelTheir own follow-up (rows below) now points away from serving: tool outputs were verified correct with zero API errors across all three looping sessions. On our stack too, the embedded template passes Phase-0 probes and scored 90.6 on tool callingSuperseded below
Follow-up forensics: the loops are exact-command repetition - one task re-ran the identical string-concat command 908 times over 2h43m, receiving a correct result every time, and still could not update its judgment (another: 500 identical runs over 1h45m)A different member of the same failure family as our stalls: ours died in a single turn at the output cap, theirs died across hundreds of turns of correct-but-ignored tool results. The completed raised-cap and thinking-off arms showed no repetition signature on our workload, but the external long-trajectory hazard remains operationally relevantNot reproduced locally
After automatic compaction the same tasks pass immediately (14/14 in 23 seconds) - long accumulated context breaks it, fresh context works; the compaction summary even misattributes the failure to "broken" tool results that were in fact correctStrong model-side signal: in-context judgment degrades as the transcript grows, and the model cannot see its own degradation. Our matrix runs sit at 64K matched context; production 3.6 has never shown this mode on the same suite. This goes in the defended verdict as an operational-hazard finding, independent of the cap artifactModel-side signal
Source 6 · MiaAI-Lab serving recipe, DGX Spark GB10 (128GB unified) + RTX 6000 PRO (96GB), vLLM NVFP4
FP8 KV cache doubles KV capacity on their rig: 1.14M to 2.30M cached tokens, 2x concurrency at 1M contextOur completed 131K A/B confirms the capacity direction: q8 KV fits one 24 GB card where f16 crashes, and q4 saves about another 2 GB. The speed direction did not transfer - true f16 was faster where it fit - so quantised KV is a capacity lever here, not a decode-speed leverCapacity direction reproduced
Working NVFP4 + FP8-KV + MTP (2 speculative tokens) recipe on vLLM, 1M context via 4x YaRNThird data point that the NVFP4 serve path is maturing (after Source 2's broken 0/69 build and Source 3's Marlin workaround), and the first showing MTP running for this model in an engine. But it is a 96-128GB memory class; the 1M-context recipe cannot fit our 24GB cards, and they publish no throughput or quality numbers at allRecipe only
No benchmarks: the repo ships configuration and capacity figures, not speed, quality, or agentic scoresIts value was operational: it showed that the MTP head survives in the NVFP4 checkpoint. Our SGLang Phase 1 run has now independently verified that path on our hardware and supplied the missing speed numbers; the source itself still contributes no benchmark resultPath independently verified
Source 7 · MiaAI-Lab serving recipe, RTX 5090 32GB (SM120 Blackwell), SGLang NVFP4
MTP speculative decoding on SM120 Blackwell: SGLang + FlashInfer using the checkpoint's built-in MTP headPhase 1 is now measured on our cards. The 1-GPU arm did not fit. On 2 GPUs, no-spec decoded at 57.32 t/s and MTP at 96.22 t/s, a +68% lift; no-spec prefill reached 22,283.5 t/s. The best arm is about 3x the candidate's llama.cpp Q4 serve, but it remains about 29% below production near 135 t/sMeasured: Qwen3.6 still wins speed
Reasoning depth is tunable per request (xhigh, medium, low - xhigh is the default) with thinking preserved across turnsThe completed dose-response settles this on our workload: thinking-on scored 35/45 at the matrix cap, raised-cap and capped-thinking each scored 39/45, and thinking-off scored 42/45. The deepest default explains the token appetite; effort-down helped, but off was bestDose response measured: off wins
Hybrid-GDN serving quirks documented: 2048-token prefill chunks (8192 causes ~600ms decode stalls), 78.4MB state per concurrent slot, no benchmarks publishedRecipe only, no verdict input. The per-slot state cost explains why this architecture concurrency-limits hard on small VRAM, and the chunking trap is worth carrying into any future engine trial on our sideNoted
Repo update (08-16): the headline now reads work in progress - does not work currently, and it adds a worked 32GB memory budget: 17-22GB weights, 78.4MB hybrid state per slot with 8 slots per request (4 lazy + 4 MTP draft), concurrency 2 via a 1.25GB mamba-cache pool, fp8 KV at 32.8KB/token, context capped at 100K, prefill chunked at 2048 (8192 chunks stall decode ~600ms on this architecture)Its implementation status remains weak external evidence, but its fit warning transferred: our 1-GPU arm did not fit. The 2-GPU Phase 1 arm did boot and measure cleanly, so the external recipe is no longer carrying the feasibility claim by itself. Phase 2 remains off because the measured MTP arm still trails production decode by about 29%Fit warning reproduced
Source 8 · Single DGX Spark (GB10), vLLM 0.27 nightly, NVFP4 + FP8 KV + MTP depth 3
First hard MTP numbers for this model: speculative decoding at depth 3 takes single-stream decode from 11.1 to 31.7 t/s (2.85x); depth 4+ crashes because the checkpoint carries only one MTP layer, and ngram speculation is ineffective for this workloadThe direction transferred but the ratio did not: our SGLang Phase 1 arm moved from 57.32 to 96.22 t/s, a 1.68x lift. That is enough to make SGLang about 3x faster than the candidate's llama.cpp Q4 serve, but not enough to catch production near 135 t/s. The source correctly predicted a useful MTP gain; our measurement closes the promotion question in production's favourDirection reproduced, smaller gain
NVFP4 beats FP8 on both throughput axes on their rig: 31.7 vs 21.9 t/s single-stream, 313 vs 225 t/s aggregate at 64-way concurrency, via a native FlashInfer kernelOur SGLang Phase 1 result confirms that NVFP4 can serve quickly on this hardware class, but it remains a speed-only result. No source has yet shown this NVFP4 checkpoint scoring cleanly on quality, so it cannot replace the quant-matched Q4 evidence in the promotion decisionSpeed path reproduced
Tool calling requires the qwen3_xml parser; the JSON-based hermes parser fails silently. First large prefill stalls ~13s on kernel JIT until warmedOur stack is unaffected - llama.cpp's embedded template with --jinja passes the Phase-0 probes and scored 90.6 on the tool suite - but a silently failing default parser is a plausible root cause for some community "3.8 cannot tool-call" reports, and both traps go in the notebook for any vLLM or SGLang trial on our sideNoted
Source 9 · Vendor benchmark table, full precision (marketing numbers, lower evidentiary standard)
Headline axis is agentic coding: Terminal Bench 73.0 vs 63.4, SWE-bench Pro 61.7 vs 53.5, QwenSWEBench 79.0 vs 49.3 against the 3.6-27B sibling, beating a frontier lab model on several rowsTheir comparator is the dense 3.6-27B, not our production 3.6-35B MoE, and they run full precision against our Q4_K_M - so the bars differ. On our suite the story is conditional: thinking-on the candidate loses to production at both caps (35 and 39 vs 41 and 42), thinking-off it wins the matched cell 42 to 41. The vendor's "large agentic upgrade" survives contact with our data only in the thinking-off configurationDifferent bar
General-knowledge gains are modest in their own table: GPQA 89.2 vs 87.8 for the 27B siblingMatches our knowledge axis, where production keeps a narrow lead at matched budget (+1.0 mean: 82 vs 80 two-card, 80-80 tie one-card; the old +4.0 included a budget-starved 74): even the vendor does not claim general knowledge as the upgrade, which is consistent with agentic-and-coding being where this checkpoint spent its training budgetConsistent
Tool-use and instruction rows are strong: IFBench 79.5 vs 69.1, LiveCodeBench 90.3 vs 83.9Direction reproduces on our deterministic tool suite (90.6 vs production's 84.4, the candidate's one proven deciding-axis win). The instruction-following claim we have not measured separately; it rides inside the agentic and quick-quality scoresPartly reproduced
Source 10 · Official vLLM serving recipe (authoritative configuration, no benchmarks)
NVFP4 weights occupy 24.6 GiB on a single Blackwell GPU at tensor-parallel 1Hard gate input for our speculative-decoding probe: 24.6 GiB of weights alone exceeds one of our 24GB cards before any KV cache, so a single-card NVFP4 serve is out - the probe would need both cards, which costs the production-parallel angle. The VRAM-fit gate the probe carries is now answered in the negative for single-cardProbe gate tightened
Official speculative-decoding config ships MTP at exactly 3 speculative tokensConverges with Source 8's measurement from the other direction: they found depth 3 optimal empirically (2.85x) and depth 4+ crashing on the single MTP layer; the vendor recipe simply never offers more than 3. The probe config is settledConverges
Thinking controls are first-class serving options: enable_thinking false, and reasoning_effort low / medium / xhigh via chat-template kwargsDirect support for our final serving recommendation: thinking-off scored 42/45 and the completed capped-thinking arm did not beat it. The official SGLang cookbook now independently names qwen3_coder as the parser for this checkpoint, resolving the earlier qwen3_xml version-drift caveat for that engineSupports our measured best config
Source 11 · llama.cpp draft-MTP A/B, same unsloth Q4_K_M GGUF, RTX 3090 24GB + RTX 5090 mobile 24GB
MTP speculative decoding works in llama.cpp itself (the July draft-mtp support) with the quantizer-preserved draft head, measuring +33% decode on a 3090 (31.0 to 41.3 t/s, acceptance 0.76-0.82) and +39% on a 5090 mobileVerified both prerequisites live on our stack within the hour: our bench binary already exposes the draft-mtp option, and our exact GGUF carries the four draft-head tensors. Their 3090 baseline of 31.0 t/s is nearly identical to our 31-33, so their +33% is a direct prediction for our cards - no engine change, no requantization, no second GPU. A same-binary MTP A/B jumped the queue - and it is now MEASURED on our cards: baseline 31.5 t/s, draft-MTP n-max 2 = 49.8 t/s, a +58% decode lift at 71.8% acceptance (3-run medians, warmup discarded, greedy 512-token completion, single slot). That beats their +33% prediction on a near-identical baseline. The May caution resolved exactly as their numbers claimed: Qwen3.6's weak draft head (0.67 acceptance, 12% slower) was the model's fault, not the engine's. Also corrects Source 4's "GGUF packs carry no MTP head": that is pack-dependent, and the unsloth pack we downloaded kept itConfirmed on our stack: +58%
Draft-depth tuning is shallow on their cards: n-max 2 wins (50.9 t/s), n-max 3 drops to 48.3 with acceptance falling 0.76 to 0.68, prose degrading faster than codeEngine-dependent optimum: the vLLM and SGLang sources converge on depth 3, llama.cpp's sweet spot is 2. Our A/B ran both and REPLICATES the ordering: n-max 2 = 49.8 t/s at 71.8% acceptance, n-max 3 = 48.8 t/s at 64.7% acceptance - depth 3 drafts more tokens (521 vs 419) but accepts a lower fraction, exactly their pattern. Their methodology also mirrors ours (3-run medians, warmup discard) and was measured at 131K resident context on quantised KV - yet another co-occurrence of the quantised-KV lever our extension arm testsReplicated: n-max 2 wins
Source 12 · CRUDbench contract task, non-thinking vs xhigh thinking (X post, @LottoLabs, 2026-08-15)
On a tightly specified find-check-delete task, non-thinking stayed literal to the contract and passed all three hidden fixtures in 27.7s with 5 model calls, 931 output tokens, and zero reasoning characters; xhigh inferred soft-deletion and status semantics from schema fields nobody asked about, failed, and retries kept reintroducing the invented conditionsSame direction as our ladder (thinking-off 42/45 vs thinking-on 35/45) and it supplies the mechanism our truncation trace could not see: the reasoning channel does not just overrun budgets, it invents requirements. The retry behavior - correcting one assumption then reinstating it - also echoes Source 5's judgment-collapse signature, where correct evidence fails to update the model's beliefs. The author's framing is careful and matches ours: not "thinking bad", but reasoning depth changes operating behavior, and extra inference only helps when the inferences are correctConsistent, adds mechanism
Their conclusion: for tightly specified tasks, non-thinking better enforces the stated contractNow a three-source convergence: their contract test, our 19-task suite, and the vendor shipping thinking-off as a first-class mode. The completed capped-thinking arm scored 39/45 against thinking-off at 42/45, so the evidence supports thinking-off as this deployment's default postureConvergent: off wins
Source 13 · community NVFP4 quantization with the MTP head preserved (vLLM, four 16GB Blackwell cards, TP4, 128K ctx, fp8 KV)
MTP at 3 speculative tokens lifts single-stream decode 49.0 to 72.6 aggregate t/s (+48%), and the gain survives concurrency: still +22% at 8-way (318.3 to 386.9)The single-stream direction reproduced twice on our stack: llama.cpp gained +58% at n-max 2, while SGLang gained +68% with EAGLE 3/1/4. The optimum remained engine-specific - llama.cpp n-max 3 slipped below n-max 2 as acceptance fell. We did not run a matched concurrency arm, so their positive 8-way result remains external evidence rather than a local claimSingle-stream gain reproduced
Their NVFP4 pack is 20.6 GB (bf16 55.6 GB compressed W4A4, with the MTP head, vision tower, and output layer kept bf16), against the official recipe's 24.6 GiB footprintPartially reopens the single-card question Source 10 closed: 20.6 GB of weights fits one of our 24GB cards, though the ~3 GB left for KV and activations is thin at any useful context. Caveat on their platform: W4A16 does not serve at all on this architecture in their vLLM version (kernel tile mismatch), so W4A4 is the only NVFP4 flavor that runs - quality of W4A4 group-16 on this model is unverified by anyone so farReopens single-card, quality unverified
Their gotchas: the bf16 draft head must stay on the quantization ignore-list or acceptance silently drops to 0% and MTP goes net-negative; the model needs a generation budget of at least 4096 or the thinking phase eats it; tool parser qwen3_xml on their vLLM versionAll three converge with our findings: the silent-degradation failure mode is the exact trap our A/B methodology (acceptance-rate logging, 3-run medians) is built to catch; the budget warning is our cap-starvation result restated by an independent operator; and the parser naming settles the Source 8 vs Source 10 drift - qwen3_xml on older vLLM, qwen3_coder on the current recipe, version-dependent not contradictoryConvergent on all three
Source 14 · vendor deployment guidance, reasoning effort and long context (relayed)
Vendor warning: in multi-turn agentic tasks, LOWER reasoning effort can raise total latency and tokens through more failures and retries (xhigh is the default effort)Our completed arm measured the opposite on this workload: capped thinking scored 39/45 and thinking-off scored 42/45 with zero truncations and zero loops, versus 35/45 with thinking on. The disagreement remains workload- and precision-scoped rather than universal, but our data governs this deploymentOur measured arm favors off
Native context 262K; YaRN static scaling extends to 1M, with the factor tuned to typical prompt length rather than maximumMatches Source 10's recipe (262K native, ~1M via overrides) - consistent, nothing new for us since our suite runs at 64K and the extension arm targets 131K resident, both inside native range with no YaRN needed. The tune-to-typical-length advice is worth keeping if a long-context serve ever happensConsistent, not load-bearing
Source 15 · Intel Arc Pro B70 inference cookbook (vLLM XPU, GPTQ-Int4 weights + BF16 MTP head, 230W card) - self-reported, provisional
MTP depth ladder at p512/g128 decode: no speculation 32.9 t/s, 1 draft token 52.0 (+58%), 2 tokens 65.8 (+100%), 4 tokens 83.7 (+154%), acceptance 93.7-100%Their +58% at one draft token is numerically identical to our llama.cpp n-max 2 measurement, from a different engine, quant format, and silicon vendor - a striking replication of the first rung. But their ladder keeps paying at deeper drafts where ours reversed (our n-max 3 lost ground as acceptance fell to 64.7%), so the depth optimum is engine-dependent, not universal: llama.cpp saturates at 2, vLLM-family engines keep gaining to 3-4 while acceptance stays above ~90%. Refines Source 11's "2 beats 3" from a model fact to an engine fact+58% replicated, depth optimum is engine-dependent
Full-context run (130,944-token prefill): decode 23.2 t/s without speculation, 56.3 with 4 draft tokens (+143% at depth)First source to show the MTP gain surviving depth. Their un-specced 23.2 t/s at full context is almost exactly our 131K figure (22.6), and they recover 2.4x of it - if that transferred to any engine our cards run, the candidate's worst axis at depth would close substantially. Single self-reported run on Intel silicon, so it stays provisional until reproducedProvisional, high value if it transfers
Source 16 · Thinking-budget sweep on an HTML canvas test (X, @KyleHessling1, 2026-08-16) - video comparison, qualitative labels, no scores
Reasoning-budget ladder for Qwen 3.8 27B: thinking OFF and 2K caps labelled broken, 6K best, 12K excellent, 24K wasteful; wall time climbs ~150s to ~550s across the sweep; claimed sweet spot 6-12K; notes the model defaults to maximum thinkingHalf confirms, half inverts our data - and the split is the finding. The "defaults to maximum thinking, responds well to caps" claim matches our dose-response ladder exactly (thinking length is a controllable cost lever on this model). But the quality ordering is the inverse of ours: our matched agentic cell peaked at thinking OFF (42/45) and every thinking dose cost 3-7 points, while his OFF is broken and quality peaks mid-budget. Reconciliation is workload, not contradiction: single-shot codegen where any mistake breaks the render rewards CoT, agentic tool-loops punish it. So the serve recommendation becomes conditional - thinking off for agentic serving (our measurement), a 6-12K cap rather than unbounded for single-shot codegen (his). Single-task, no repeats, no scoresDomain-split: confirms the cap lever, inverts the OFF verdict
Source 17 · Practitioner report on thinking budget and role-splitting (X, @TheAhmadOsman, 2026-08-16) - anecdote, no numbers
Claims Qwen 3.8 27B "improves tremendously" at a medium thinking budget, and that his preferred setup is another model doing the thinking and planning while 27B runs implementation with no thinking at allThe second half is our agentic finding restated as a user preference: when planning lives outside the model - a planner model, or an agent harness driving a tool loop - the executor runs best thinking-off, which is exactly the cell where the candidate scored its 42/45. The first half lands beside Source 16's mid-budget sweet spot for standalone use. Three independent lines of evidence (our measured ladder, Source 16's sweep, this report) now triangulate one picture: thinking budget is a real quality lever on this model, the default is too high, and OFF is correct when the model is the executor rather than the plannerAnecdote: corroborates the workload split
Source 18 · Official SGLang Qwen3.8-27B cookbook, H200 FP8 balanced cell and Blackwell NVFP4 guidance
The H200 FP8 balanced cell is a verified single-node deployment recipe: FP8 weights, FlashInfer, 32,768-token prefill chunks, 0.85 static-memory fraction, and the Qwen reasoning and tool parsers. The page publishes no speed result for that cellNot a head-to-head with our 2x24GB NVFP4 probe. H200 FP8 uses a different checkpoint precision, memory class, kernel path, and prefill policy. It cannot replace our measured 57.32 / 96.22 t/s decode result or production's measured ~135 t/s. Winner on our deployed hardware remains Qwen3.6Recipe only: Qwen3.6 still wins measured speed
The official MTP overlay is EAGLE with 3 steps, top-k 1, and 4 draft tokens; small Blackwell guidance uses 2,048-token prefill chunks and warns that the default Mamba memory ratio can clamp concurrencyOur Phase 1 script already used the same EAGLE 3/1/4 flags, FP8 KV, and 2,048-token chunks, so the official page exposes no missed single-stream decode lever. We did not pin the Mamba ratio, but that governs state-pool concurrency rather than the one-request decode probe used for the promotion gate. No rerun or Phase 2 is justifiedCore recipe independently matched

The matched-budget repeats are done and they settle the quality axis cleanly. The composite is a config split: production +0.4 on one card (92.1 vs 91.7), candidate +1.2 on two (93.8 vs 92.6), candidate +0.4 on the mean - inside the ~2.6 noise band. The candidate's old tight-cap 93.8 was partly a budget artifact and its 89.8 entirely so. What survives matching: maths is a real candidate win (+5.2 mean, 97/100 vs 93.3/93.3), knowledge narrows to production +1.0, instruction-following goes to production +2.5, and token efficiency flips - production is ~16% leaner per scored item once caps match. The MTP opportunity flagged here is no longer hypothetical: measured on our own engine at +58% decode (see Source 11).

Campaign plan

Standard intake chain first, then the extension arms that turn a screening result into a defended verdict.

Final verdict: keep production Qwen3.6-35B - Qwen3.8-27B stays lab-only as the strongest quality candidate ever benched here

The campaign is closed and the call is made. Qwen3.8-27B Q4_K_M fails all three promotion clauses: the quick-quality composite is a config split whose +0.4 mean sits inside the ~2.6 repeat-noise band (not a beat), the speed-weighted agentic /100 reads 71.1 vs 83.5 in production's favour, and the Q4 throughput gap - 3.4 to 4.1x slower at roughly 5x the energy per token - is a material speed regression by any reading of the bar. SGLang Phase 1 materially narrows that axis: NVFP4 MTP reaches 96.22 t/s, about 3x the llama.cpp Q4 serve, but it remains about 29% below production near 135 t/s and does not flip the speed clause. No clause is met; the bar requires all three. Production Qwen3.6-35B stays and SGLang Phase 2 stays off.

What makes this a lab-keeper rather than a discard: it is the first candidate to win a matched agentic cell against production (42 vs 41, thinking off, inside noise but at the measured suite ceiling - Opus 4.6 lands ≈41 on the same tasks with production's exact tier profile), it holds the best expert-tier score on file (16/18), the best tool-calling score on file (90.6 vs 84.4, +6.2 proven), a proven maths win (+5.2), and a working MTP speculative-decoding path worth +58% decode on our own engine. Operational flags if it is ever served: thinking must be OFF (the dose-response ladder 35/39/39/42 shows its own reasoning channel is the deficit), TC-SF-01 fired a destructive delete without confirming, and production stays ~16% token-leaner per scored item.

The campaign also paid for itself in methodology: card-bench has silently served q8_0 KV cache since inception (every published number is a q8-KV number - now a documented, measured default: true f16 KV is +10% decode / +21% prefill at 131K but cannot fit 131K on one 24 GB card), the kvq8 arm doubled as an exact-repro check and passed within 0.2 t/s, and the Q8_0 weight-quant arm confirmed Q4_K_M's discount is speed-positive (+56% decode at 8K) and ~10.5 GB lighter, not quality-funded. Production was restored at campaign end and verified live: both services active, health endpoint returning 200.

The agentic axis now has a three-configuration ladder, and the story turned. Thinking on at the matrix budget the candidate loses clearly: 35/45 vs production's 41/45, with every stall traced to hard truncation at the shared 4096-token cap. Raising the cap to 16K recovers half the gap (39/45 vs production's 42/45 at the same cap) but converts the worst task's truncation deaths into wall-clock timeout deaths - the thinking-token appetite is a real operational cost either way. Then the surprise: with thinking off entirely, at the original tight cap, the candidate scored 42/45 - zero truncations, zero loops, tying the best production number on file and beating the thinking-on matrix comparator. The agentic deficit is not "cannot do agentic work"; it is the cost of this model's own reasoning verbosity, and it vanishes when the reasoning channel is closed. That honesty flag is now closed: a production thinking-off cell was measured in the identical configuration and scored 41/45 - production gains nothing from turning thinking off (its fast MoE decode was never starved by the reasoning channel), so the matched cell reads candidate 42, production 41: the candidate's first outright win on the agentic axis. Two qualifiers stay attached: a one-point edge sits inside the repeat-noise band, and production's best number anywhere on file remains 42 (thinking on, 16K cap). The completed ladder's summary line: thinking costs this candidate 3 to 7 points; it costs production nothing. The matched-budget quick-quality repeats then settled the last caveated axes: the composite is a config split (production +0.4 on one card, candidate +1.2 on two, candidate +0.4 on the mean - inside noise), maths is a proven candidate win at +5.2, knowledge narrows to production +1.0, instruction-following goes to production +2.5, and token efficiency flips to production (~16% leaner per scored item; the earlier candidate lead was an output-cap artifact). Where that leaves the ledger: Qwen3.8 holds tool calling (+6.2, proven), maths (+5.2, proven), refusal behavior, VRAM footprint, and the matched agentic cell (42 vs 41, inside noise); Qwen3.6 holds every speed, power, and token-efficiency axis (proven, 3.4-4.4x on throughput) plus narrow leads on knowledge and instruction-following. The promotion bar is still not met: the agentic /100, which weighs speed, reads 71.1 vs 83.5 in production's favour. A frontier calibration pass reframes what the /45 numbers mean: Opus 4.6, driven through a first-party CLI harness over the same tasks, sandbox, and oracle, passed 17/19 tasks and failed exactly the two expert tasks production drops - at the same 1/3 rate on 3-run repeats. Full 3-run cells now exist for every expert and hard task, and they firm the picture: Opus lands expert 14/18 and hard 12/12 - production's tier profile to the point, one under the candidate's matched-cell 42. The suite's effective ceiling is 41-42/45, a frontier model sits on it alongside both finalists, and the /45 axis is saturated; what still separates the two local models is speed, power, and token efficiency, all of which production holds.

Opus 4.6 against both Qwens

Opus 4.6 beats both local models outright on exactly two axes, and both are purple-marked in the head-to-head table: the quick-quality composite (94.4 vs the candidate's 92.75 and production's 92.35 config means) and knowledge (86.0 vs 80.0 and 81.0 - the axis where frontier-scale pretraining and its contamination exposure show up most directly). Everywhere else at least one Qwen matches or beats it. The candidate's maths mean of 98.5 edges Opus's 96.7. Instruction-following is a 95.0 tie with production. The long-context needle is a three-way 100. The refusal rate ties the candidate's best cell at 0.00. On the agentic suite Opus lands ≈41/45 - production's exact tier profile, one point under the candidate's matched-cell 42, on a suite whose measured ceiling is 41-42, so the /45 axis separates nobody at the top. Answered accuracy (0.36) is the one Opus number below the locals, and it stays flagged rather than concluded because the harness had no temperature control and a different system prompt. The overall read: at this suite's difficulty, a frontier model separates from well-served local 27-35B weights on knowledge breadth and a composite it drags up, and on nothing else that survives the noise band - while the locals run on two workstation cards at zero marginal cost per token.

Promotion bar: beat production on quick-quality composite AND agentic /100, quant-matched, without material speed or VRAM regression. Whether a speed regression is "material" is the operator's call, weighed against how large a quality win actually lands.

Update log

08-16Post-close AtomicChat quant diagnostic completed. Its imatrix AD-Q4_K tied unsloth Q4_K_M on plain decode (31.20 vs 31.21 t/s), edged the same MTP n-max 2 path (46.79 vs 46.31 t/s, 66.4% vs 64.3% acceptance), and added only 52 MiB VRAM. On the strict matched 16K-output quick-quality cell it scored 93.9 vs 91.7: knowledge 84 vs 80, instruction 95 vs 90, maths 96.7 tied, needle 100 tied, zero errors and zero truncation. AtomicChat wins this single matched quality run, but the +2.2 composite remains inside the known same-cell noise band and it lacks the full tool and agentic evidence attached to the unsloth campaign artifact. The qwen38 switcher therefore stays on unsloth; the keep-production verdict is unchanged.
08-16Source 18 filed: the official SGLang Qwen3.8-27B cookbook. Its linked H200 FP8 balanced cell is authoritative deployment guidance, not a benchmark, and is not hardware- or quant-matched to our 2x24GB NVFP4 Phase 1. The transferable controls were already present in our probe: EAGLE MTP 3/1/4, FP8 KV, and the small-Blackwell 2,048-token prefill chunk. The page's Mamba-ratio warning concerns concurrency capacity, not our single-request decode gate. It adds confidence that the probe used the intended serve path but supplies no missing optimization that could close the remaining ~29% production gap. Verdict unchanged; Phase 2 stays off.
08-16SGLang Phase 1 folded into the durable snapshot. The 1-GPU NVFP4 arm did not fit. On 2 GPUs, no-spec measured 57.32 t/s decode and 22,283.5 t/s prefill; MTP raised decode to 96.22 t/s, +68%. That is about 3x the candidate's llama.cpp Q4 serve, but still about 29% below production near 135 t/s. The speed clause does not flip, Phase 2 stays off, and the keep-production verdict is unchanged.
08-16Source 16 filed: a thinking-budget sweep of the candidate on an HTML canvas test (X, @KyleHessling1). His ladder - OFF/2K broken, 6K best, 12K excellent, 24K wasteful - inverts our agentic dose-response (thinking OFF won our matched cell at 42/45) while independently confirming the mechanism behind it: the model defaults to maximum thinking and responds cleanly to caps. Read as workload-conditional, not contradictory: CoT helps single-shot codegen where any mistake breaks the render, and hurts tool-loop agents. The serve-thinking-off recommendation stands for agentic serving; for single-shot codegen his 6-12K cap beats unbounded. Source 17 (a second practitioner, independently, same day) closes the loop: medium budget standalone, no-thinking execution under an external planner - the exact split our ladder measured. The thinking-budget story is now triangulated from three independent directions.
08-16Opus 4.6 quick-quality COMPLETE, and the "would need API credits" caveat is retired: the same first-party claude -p path that ran the agentic column drove the full frozen suite (a driver swaps the suite's chat call for a claude -p subprocess; smoke gate passed first, 703.8s wall, concurrency 4). The numbers, flags carried: knowledge 86.0 - the best on file, with locals at 80-82; maths 96.7 with zero errors (the candidate's 97/100 keeps that axis by 1.8 on the mean); instruction-following 95.0, tying production; refusal rate 0.00; and a corrected composite of 94.4, the best composite on file. The correction is itself the most instructive finding of the run. The first pass printed composite 69.4 because the needle category scored 0/6 - not because a frontier model cannot find a needle, but because Claude Code treats its output cap as a hard ERROR rather than a truncation point, so the suite's 64-token needle cap killed four runs outright. Of the two answers that did return, one was harness identity confusion at 131K depth, and the other was the best qualitative result of the night: Opus found the buried needle and refused to relay it, flagging it as a suspected prompt injection. With the cap floored at 512 the needle rerun scored 100.0 on six exact extractions. Standing flags on every Opus number: harness system prompt injected, no temperature or seed control, tools disabled, knowledge and maths contamination-exposed, needle cap 512 vs the locals' 64. Also this update: Source 15 (an Intel Arc Pro B70 cookbook) independently lands +58% at one MTP draft token - the exact number our llama.cpp A/B measured - while its ladder keeps climbing to +154% at four drafts, making the draft-depth optimum an engine fact rather than a model fact; and Source 7's RTX 5090 SGLang recipe now carries a work-in-progress banner plus a worked 32GB memory calculator, adopted into the probe planning.
08-16Post-close addendum on request: the Opus 4.6 per-task detail is now published as its own section rather than compressed into one column cell. Every expert and hard task shows its 3-run cell beside the two local models' matched thinking-off cells, making the reportable structure visible at a glance: expert_money is the candidate's outright differentiator (3/3 where production and Opus both sit at 1/3), expert_calc is the task nobody clears, and everything else on those tiers is a three-way clean sweep except hard_base, where the candidate gives one back. No numbers changed; the verdict stands.
08-15 23:55CAMPAIGN CLOSED. The Q8_0 weight-quant arm finished clean (sha256-verified 27.1 GiB download, chained card-bench, exit 0) and settled the last open axis: Q4_K_M beats its own 8-bit sibling on decode by +56% at 8K (31.9 vs 20.4 t/s) and +40% at 131K (22.6 vs 16.2), prefills +9% faster at depth, and saves ~10.5 GB VRAM - on this architecture the 4-bit discount is pure win, and the campaign's quality numbers (42/45 at the suite ceiling, on Q4 weights) confirm it is not quality-funded. The SGLang NVFP4+MTP probe is formally skipped: at close time the working belief was that this GPU architecture had no SGLang backends, and the same-engine MTP A/B had already delivered the number. [Superseded 08-16: that backend belief was wrong. Phase 1 later ran clean on these cards and measured 57.32 / 96.22 t/s; see the 08-16 entries above. The skip reason changed, the verdict did not.] Final verdict written into the repo's decision log and onto this page: keep production Qwen3.6-35B - the candidate fails all three promotion clauses (composite inside noise, agentic /100 71.1 vs 83.5, 3.4-4.1x speed regression at ~5x energy per token) - but Qwen3.8-27B stays as the strongest quality candidate the lab has ever benched, with a matched-cell agentic win at the measured ceiling, the best expert tier and tool-calling scores on file, a proven maths win, and a +58% MTP path if it is ever served (thinking OFF, with the destructive-confirm tool-calling miss on record). The methodology dividend is permanent: the q8-KV silent default is now documented and measured, and the exact-repro check passed within 0.2 t/s. Production was restored at campaign end and verified: both services active, health endpoint 200. This page is final.
08-15 22:50KV-precision A/B closed, and the f16 baseline refused the tidy ending: it is FASTER at depth, not equal. True-f16 KV decodes 24.9 t/s at 131K on 2 GPUs against q8's 22.6 (+10%) and prefills 49.7 vs 41.2 (+21%), so the lab's silent q8_0 default has been costing real depth speed on every published card-bench number - the opposite direction to Source 3's claim that quantising KV recovers depth throughput. The catch that keeps q8 as the right default anyway: f16 KV cannot fit 131K on a single 24 GB card (the 1-GPU cell segfaulted; the 2-GPU peak was 25.9 GiB), so on this hardware quantised KV is a fit requirement at depth, and at 8K all three KV types decode within ~1% of each other. Depth's main tax is attention compute either way: even f16 slides 32.1 to 24.9 from 8K to 131K. The blocked kvq4 1-GPU cell also refilled clean (20.9 t/s, 19.0 GiB - q8 is the 1-GPU speed pick, q4 the VRAM pick). Next arm is already moving: the Q8_0 weight quant (27.1 GiB, 2-GPU only) is downloading with a chained card-bench behind it.
08-15 21:55The quantised-KV arm found something better than what it was looking for: a methodology fact. The kvq8 cells came back IDENTICAL to the parent rows - decode within 0.2 t/s and VRAM identical to the MiB - and the server logs explained why: card-bench passes --cache-type q8_0 by DEFAULT on every llama row, so every card-bench number this campaign (and this lab) has ever published was already a quantised-KV number, and the kvq8 row duplicated the parent config exactly. Two consequences. First, the accidental duplicate is a gift: it is a clean exact-repro check of the parent row, and the methodology passes it - 21.17 vs 21.19 t/s at 131K on separate runs half a day apart. Second, Source 3's depth-recovery claim inverts: the 30.6-to-21.2 depth slide happens WITH q8 KV, and dropping to q4_0 KV recovers nothing further (131K decode 22.26 vs 22.58 t/s on two cards, ~2 GB VRAM saved) - so on this stack, long-context decode loss is not KV-precision-bound, unlike Source 3's vLLM FP8-KV setup. A true-f16 baseline row (kvf16, last-value-wins override) is chained behind the running cells to close the A/B from the other side: if f16 also decodes ~21-22.6 at 131K, KV precision is confirmed irrelevant to depth throughput on this architecture.
08-15 21:20Opus 4.6 tier repeats COMPLETE - a perfect sweep. All 24 runs passed (the four remaining expert tasks and all four hard tasks, 3/3 each, walls 12-34 seconds), so the frontier reference now carries full 3-run cells on every expert and hard task. Firmed tallies: expert 14/18 (only expert_calc 1/3 and expert_money 1/3 miss - production's exact failure profile), hard 12/12 (also production's exact number), ≈41/45 overall. The reportable comparison: a frontier model driven over the same tasks, sandbox, and oracles lands on production's tier split exactly, and one point under the candidate's thinking-off 42. The candidate's expert 16/18 is the best expert tier ever recorded on this suite - it is the only model of the three to pass expert_money at better than 1/3 - while its hard 11/12 gives one back. The quantised-KV card-bench arm chained on schedule at 21:12:31, first cell kvq8 / 8K / 1 GPU.
08-15 21:06Two arms launched as one chained job. First, on request: Opus 4.6 repeats on the high-value tiers - 3-run cells for the four expert tasks and four hard tasks that only had single first-pass runs, which will give Opus a full 3-run expert /18 and hard /12 directly comparable to the local finalists' tier splits (expert_calc and expert_money already have their x3 cells). First cell passed 21 seconds in. Second, chained to start only after the repeats finish so Docker CPU load never overlaps the 10ms power sampling: the quantised-KV card-bench arm - q8_0 and q4_0 KV-cache rows of the same candidate GGUF at 8K and 131K contexts on both GPU configs, testing Source 3's claim that quantised KV recovers decode throughput at depth (f16-KV baseline: 21.2/22.6 t/s at 131K).
08-15 21:02Two page changes on request. The head-to-head table no longer scrolls sideways - the long per-cell notes were inheriting nowrap and stretching the table off-screen; they now wrap. And Opus 4.6 gets its own reference column in the head-to-head: real values on the agentic rows where it was measured, explicit not-run / n-a everywhere else, in its own violet so nobody mistakes the frontier reference for a third contender. Quantised-KV diagnostic arm (q8_0 and q4_0 KV at 8K + 131K context, both GPU configs) is being registered and launched next.
08-15 20:49Opus 4.6 calibration COMPLETE, and it lands the most clarifying single result of the campaign. First pass: 17/19 tasks, total wall 380 seconds, the only failures being expert_calc and expert_money - the exact two tasks where production loses its points. Immediate 3-run repeats of those two: expert_calc 1/3, expert_money 1/3 - the same tasks, at the same rate, as production's cells (both 1/3) and matching the candidate's expert_calc 1/3. Read: a frontier model cannot clear those two tasks either, so the suite's effective ceiling is about 41-42/45, and both production (41) and the thinking-off candidate (42) are already sitting on it. The /45 axis is saturated and stops discriminating at the top; the axes that still separate the finalists are speed, power, and token efficiency. Flag carried: different harness, so the /45 is comparable but turn and token metrics are not, and single-run task passes were not all repeated - the ceiling figure is the two repeated failure cells plus 17 first-pass passes. GPU arms un-held: quantised-KV 131K is next on the cards.
08-15 20:35Opus 4.6 calibration arm built and running. The sanctioned path: the first-party Claude Code CLI (claude -p on the Max subscription) drives the same 19 agentic tasks inside the same sandbox, oracle, and workspace contract as every local arm - a new driver script slots into the harness via a per-model entrypoint, with the OAuth token passed into the container as an environment name only, never on a command line or disk. Smoke test passed first try; the full 19-run pass launched 20:35 and is moving at roughly 20-45 seconds per task. The row lands flagged: a different harness means the /45 is comparable but turn-count and token-efficiency metrics are not. Purpose is calibration, not competition - it tells us where the suite's ceiling sits for a frontier model, which prices what 41-42/45 from a local 24GB-card model actually means. GPU arms (quantised KV at 131K, then Q8_0 two-card) are held until this pass finishes so nothing perturbs a future speed window.
08-15 20:31Matched-budget quick-quality repeats COMPLETE, and they rewrite three rows. With the candidate re-run at production's 16K reasoning budget: composite 91.7 (1 GPU) and 93.8 (2 GPU) vs production's 92.1 / 92.6 - a config split, candidate +0.4 on the mean, inside the ~2.6 noise band, so the composite axis ends as a wash with a candidate lean. The real separations: maths goes to the candidate at +5.2 (97/100 vs 93.3/93.3) - a proven win, not noise; knowledge narrows from the old +4.0 to production +1.0 (the 74 that drove the old gap was budget starvation); instruction-following goes to production +2.5. And the token-efficiency story flips: at matched caps the candidate spends 766/906 tokens per scored item vs production's 674/725 - production is ~16% leaner, and the earlier "candidate leads 568 vs 674" row was an artifact of comparing a capped run against an uncapped one. Source 1's 2.9x token-appetite claim now reproduces in direction on our stack. The old budget-2048 rows are archived, not deleted.
08-15 19:00MTP speculative decoding MEASURED on our own engine and GGUF, and it delivers more than any source predicted: baseline 31.5 t/s decode, draft-MTP n-max 2 = 49.8 t/s (+58%) at 71.8% acceptance, n-max 3 = 48.8 t/s (+55%) at 64.7% acceptance. Source 11's llama.cpp ordering replicates exactly - depth 2 beats depth 3 because deeper drafts accept a lower fraction - and the May-2026 Qwen3.6 result (12% slower, 0.67 acceptance) is now confirmed as the model's weak draft head, not the engine. Caveat kept honest: this is one greedy technical-prose prompt, the friendliest terrain for speculation - agentic and code traffic will accept less, so +58% is the upper bound, not the expectation. Even so, this moves the candidate's worst axis: the decode gap vs production narrows from 4.4x to roughly 2.8x with zero quality cost (speculative decoding is lossless). Four MTP confirmations across four configs now, ours being the only one on the exact production-candidate serving path.
08-15 19:35The missing cell is measured and the tie broke in the candidate's favour. A production thinking-off run in the identical configuration (thinking off, matrix 4096 cap, 64K ctx) scored 41/45 in 26 minutes - misses expert_calc 1/3 and expert_money 1/3, hard tier a clean 12/12. So production does NOT gain from turning thinking off: its score with thinking on at this cap was also 41, because its fast MoE decode never gets starved by the reasoning channel. The matched-cell headline becomes candidate 42, production 41 - the candidate's first outright agentic win, with two qualifiers carried: one point is inside the repeat-noise band, and production's best anywhere on file remains 42 at the 16K cap. The completed four-arm ladder now reads: thinking costs the candidate 3 to 7 points and costs production nothing. Also this hour: check-eval-coverage GATE PASS for the candidate (card-bench all declared GPU configs + agentic both present). Remaining before the verdict: matched-cap quick-quality repeats, quantised-KV 131K arm, Q8_0 two-card arm.
08-15 18:52Capped-thinking arm FINAL: 39/45, and the reasoning ladder is complete - thinking-on 35, raised-cap 39, capped-2048 39, thinking-off 42. The answer to "does a little thinking beat none" is NO on this suite: any thinking at all costs about 3 points against thinking-off, and the failure modes differ by dose. At 2048 the expert block held up well (15/18) but hard_base was wiped 0/3 by clean 630-second wall-clock timeouts mid-work - 14-17 productive turns, no loops, no truncation - the per-turn thinking overhead simply starves the time budget on the longest-horizon task, the one place agentic work most resembles production use. This also arbitrates Source 14: the vendor's warning that lower effort raises latency and tokens through failures and retries did not materialise anywhere in the ladder - lower effort was monotonically equal-or-better, and the failure signature of MORE thinking was timeouts, not retries. Serve-thinking-off stands as the recommendation, now with a four-point dose-response curve behind it. Next up: the llama.cpp draft-MTP decode A/B (n-max 2 vs 3) on the freed GPU.
08-15 16:33Source 12 upgraded from relayed observation to cited primary source: the CRUDbench comparison is a public X post (@LottoLabs, 2026-08-15) with hard numbers - non-thinking passed all three hidden fixtures in 27.7 seconds, 5 model calls, 931 output tokens, zero reasoning characters, while xhigh burned far more compute repeatedly reasoning itself away from the simple correct solution. The author's own framing lands where our ladder does: reasoning depth changes operating behavior, and for tightly specified tasks non-thinking acts as a constraint against inventing requirements. No verdict change - the row already said this - but the evidence class moved from second-hand to quotable.
08-15 16:25Source 14 (relayed vendor deployment guidance) is the first source to run AGAINST the thinking-off finding: the vendor warns that lower reasoning effort in multi-turn agentic tasks can raise total latency and tokens through more failures and retries, and ships xhigh as the default. Our suite measured the opposite (thinking-off 42/45, zero retries-by-failure signatures, vs thinking-on 35/45), and Source 12's contract test sides with us - but the honest read is that the disagreement is now on the record and the capped-thinking arm currently running is the exact arbiter between the two positions. The long-context half of the guidance (262K native, YaRN to 1M tuned to typical length) matches the official recipe and touches nothing we run today.
08-15 16:22Source 13 (a community NVFP4 pack with the MTP draft head preserved in bf16) makes it three independent MTP confirmations across three serving configs. Their numbers: +48% single-stream at 3 speculative tokens, and - the real news - still +22% at 8-way concurrency, where our May measurement on Qwen3.6 collapsed to -75%. If that concurrency behavior transfers, MTP on 3.8 is not just a single-stream trick. Their 20.6 GB footprint also partially reopens the single-card NVFP4 question the official recipe's 24.6 GiB had closed, though with ~3 GB left for KV it would be thin, and nobody has verified W4A4 group-16 quality on this model. Their gotchas independently restate two of our findings: a silently-broken draft head shows as 0% acceptance and net-negative speed (exactly what acceptance-rate logging in our queued A/B catches), and the thinking phase eats any budget under 4096 (our cap-starvation result, from a stranger's serving notes).
08-15 16:16Two more sources, and one of them unblocks the decode story on our own engine. Source 11 measures MTP speculative decoding in llama.cpp itself - same unsloth Q4_K_M GGUF we run - at +33% decode on a 24GB 3090 whose baseline (31.0 t/s) is nearly identical to ours, acceptance 0.76-0.82 at draft depth 2. Verified immediately on our stack: the bench binary exposes the draft-mtp option and our GGUF carries the four draft-head tensors, so a same-binary A/B jumps the queue ahead of the SGLang probe (which is now deprioritised). Honesty note carried on the row: the identical engine feature measured 12% slower on Qwen3.6-27B in May at 0.67 acceptance - the gain rides on 3.8's better-trained draft head, which is exactly what the A/B will confirm or refute. Source 12 is a relayed CRUDbench contract test: non-thinking passed by staying literal, xhigh thinking invented soft-delete semantics nobody asked for and kept reinstating them through retries. That supplies the mechanism behind our thinking-off 42/45: the reasoning channel does not just overrun budgets, it invents requirements. Serve-thinking-off is now a three-source convergence (their contract test, our suite, the vendor's first-class flag).
08-15 16:07Tenth source: the official vLLM serving recipe. Three consequences. First, the NVFP4 weight footprint is 24.6 GiB at tensor-parallel 1 - more than one of our 24GB cards holds before KV, so the speculative-decoding probe is single-card-infeasible and becomes a 2-GPU question. Second, the official MTP config ships exactly 3 speculative tokens, converging with Source 8's empirical depth-3 optimum - the probe config is settled. Third, and best for the verdict: thinking-off (enable_thinking false) and a low/medium/xhigh effort ladder are first-class vendor serving options, so the configuration that just scored 42/45 is a shipped mode, not a lab workaround. Capped-thinking arm meanwhile through its first task: expert_calc 1/3 - now 1/3-or-worse in all four configurations, making it look task-hard rather than mode-sensitive.
08-15 15:58Reasoning-off arm FINAL: 42/45 at the matrix 4096 cap - the candidate's best agentic number, tying production's best on file (42/45 raised-cap) and beating the thinking-on matrix comparator (41/45). Zero truncations, zero loops, zero timeouts except one: the three misses are expert_calc 1/3 (short genuine attempts, 3-5 turns - this task has never gone above 1/3 in any configuration) and one hard_base timeout already screened clean of repetition. Thinking-off cured both failure modes at no cost anywhere else, which reframes the agentic deficit as reasoning-verbosity overhead rather than capability. Caveat carried into the verdict: production's comparators are thinking-on runs; no production thinking-off cell exists. Capped-thinking arm (2048 budget) launched 15:54 and verified on the server args - does a little thinking beat none? Ninth source also added: the vendor's own benchmark table, which claims agentic coding as the headline upgrade vs the 27B sibling; on our data that claim holds only in the thinking-off configuration, and their comparator is not our production model.
08-15 15:36Eighth source is the first to publish hard MTP numbers for this model: on a single DGX Spark under a vLLM 0.27 nightly, MTP depth 3 lifts single-stream decode from 11.1 to 31.7 t/s - 2.85x, with depth 4+ crashing on the checkpoint's single MTP layer. That quantifies the decode-recovery thread: a transferring ratio of that size would cut the candidate's 4.4x decode deficit to roughly 1.5x, so the queued SGLang MTP probe now has a number attached. Also carried: NVFP4 outruns FP8 on both throughput axes on their rig (its quality remains unverified everywhere), and a silent tool-calling trap - the default JSON parser fails silently and the model needs the qwen3_xml parser; our llama.cpp template path is unaffected. Meanwhile the reasoning-off arm reached 37/40 with hard_base recovered to 2/3 and mediums 9/9 - thinking-off has already banked more passes than the thinking-on matrix run's 35/45 with five runs still to record.
08-15 15:22Reasoning-off arm at 17/20 - the strongest run of the three at this point (thinking-on matrix was 13/20, raised-cap 15/20 through the same tasks). expert_money is 3/3 for the first time, and hard_base produced its first clean pass since the matrix run: 7 turns, 135s, no cap pressure at all. Its r2 still timed out at 27 turns, but the transcript screens clean - max consecutive identical command is 2, worst any-position repeat is 3, ordinary trace-debug retries rather than Source 5's hundreds-of-identical-runs loop. Early read: turning thinking off costs nothing on the expert tier and removes the failure modes, directly contradicting the thinking-off prior from Source 5's stack.
08-15 14:25Raised-cap arm final: 39/45 vs production's 42/45 at the same 16K cap. The output cap explains half the matrix gap (35 to 39 of the 6-point deficit) but not all of it: production still wins by 3 with the artifact removed. What recovered: csv_dialect 0 to 3, calc 0 to 1, duration and lru back to 3/3, all mediums and smalls clean. What got worse: hard_base 2/3 to 0/3, all three runs timing out at the 630s wall mid-debug with zero repetition loops - the extra thinking room converts truncation deaths into timeout deaths on the one task that needs long trajectories. Reasoning-off arm launched at the matrix 4096 cap (verified --reasoning off on the server args): if thinking-off holds the score without the stalls, the operational picture changes; Source 5's thinking-off loops are the prior against.
08-15 14:18Seventh source, same author, RTX 5090: the first recipe in our exact GPU architecture class (SM120 Blackwell) with MTP speculative decoding actually running - SGLang + FlashInfer + EAGLE off the checkpoint's own MTP head. Decode is the candidate's worst axis, so a feasibility probe joins the queue, gated on NVFP4 quality verification and whether the checkpoint plus KV fits a 24GB card. It also confirms per-request tunable reasoning depth with xhigh as the family default, a direct prior for the effort-level arms. Still no published benchmarks from this author.
08-15 14:10Sixth source added: a DGX Spark / RTX 6000 PRO serving recipe for 3.8-27B (vLLM, NVFP4, FP8 KV, MTP, 1M context via YaRN). No benchmark numbers, so no verdict input, but two operational reads: FP8 KV doubles cached-token capacity (independent confirmation of our queued quantised-KV arm from the capacity side), and the NVFP4 checkpoint retains a working MTP head, keeping decode recovery alive as a future path. Their memory class (96-128GB) is out of reach of our cards.
08-15 13:59Hard tier complete at the raised cap: hard_duration and hard_lru both recovered to 3/3 (were 2/3), hard_intervals held 3/3, but hard_base finished 0/3 - all three runs timed out at the 630s wall with heavy genuine tool use (13-16 turns, 15-20 tools, hexdump/import-trace debugging, zero repetition in all three transcripts). hard_base is now the one task where the raised cap made things strictly worse than the matrix run's 2/3. Running total 26/32; medium tier under way and passing fast (85s/run).
08-15 12:57Raised-cap arm at 15/20 and the picture is now two-sided. Expert tier finished 15/18 vs 11/18 at the matrix cap - the truncation diagnosis holds for csv_dialect (0/3 to 3/3) and calc (0/3 to 1/3). But hard_base, which scored 2/3 at the matrix cap, is 0/2 here: both runs hit the 630s per-run wall clock mid-work (16 and 12 turns of real debugging - hexdumps, import tracing, no repetition loop in either transcript). The raised cap trades truncation failures for timeout failures on this task: more thinking budget per turn means fewer tasks fit the clock. That trade-off is itself a finding against the model's thinking-token appetite.
08-15 11:47Raised-cap arm early returns confirm the truncation diagnosis: expert_csv_dialect, 0/3 at the matrix cap, is 3/3 at 16K; expert_calc, also 0/3, has its first pass. Every run so far engages cleanly with multi-turn tool use (5 to 15 turns), zero cap truncations, and no sign of Source 5's repetition-loop signature. 8/10 rows passed through the first four expert tasks.
08-15 11:22Source 5 follow-up: their forensic trace shows the agent loops are exact-command repetition (908 identical runs over 2h43m on one task) with correct tool outputs and zero API errors throughout - and the same tasks pass in seconds after compaction wipes the context. That exonerates the template/engine and indicts the model's long-context judgment. Two consequences here: the raised-cap arm's run logs will be screened for the same repetition signature, and long-context judgment collapse joins the defended verdict as its own finding, separate from the output-cap artifact.
08-15 11:07Fifth source: a ts-bench coding-agent run (UD_Q4_K_XL, thinking off, llama.cpp) reports 3.8-27B significantly worse than 3.5/3.6 with infinite agent loops - the same failure signature our harness logged before the cap-truncation trace. Notable: their loops occur with thinking OFF, which is a prior against our queued reasoning-off arm curing the stalls on its own. They suspect a template or llama.cpp defect; our raised-cap and reasoning arms will separate model from serving config.
08-15 10:52Standard chain complete (10:28, all eleven steps ok). Final agentic: 35/45, official 71.1/100, vs production's 41/45, 83.5/100. All six lost rows are hard/expert tier: expert_calc 0/3 vs 1/3, expert_csv_dialect 0/3 vs 2/3, and one dropped run each on hard_base, hard_duration, hard_lru; the other 14 tasks matched production, including 3/3 sweeps of debug_eventbus, reservations and scheduler. The runner auto-restored production and verified health, then the stack came back down for the extension arms, which run back-to-back with one final restore at campaign end. The raised-16K-cap diagnostic arm launched at 10:48 and is mid-run; production's own raised-cap comparator is already on file at 42/45.
08-15 10:15Stall root cause found: every zero-tool-call failure is a hard truncation at the suite's shared 4096-token per-response output budget (finish reason "length", exactly 4096 tokens out, every time). The thinking-native candidate reasons past the budget before acting; production averages 3216 tokens and never touches it. Context is not the issue (already 64K, matched). A raised-cap diagnostic arm is queued as the first extension run, interlocked with the reasoning-level arms (full / capped / off), and will be reported separately from the matrix number per suite policy. One failure had clean tool use and no truncation, so the cap does not explain everything.
08-15 10:00Agentic axis decided mid-run: 18/27 through nine tasks means even a perfect finish caps at 36/45 vs production's 41/45 at matched 64K context. Promotion bar can no longer be met. Failure mode is concentrated: two tasks went 0/3 by generating ~140s without a single tool call (turns=1, tools=0); five other tasks went 3/3 or 2/3 with clean multi-step tool use. Verdict panel updated to leaning-no.
08-15 09:35Fourth source: llama-bench numbers for the same GGUF quants on Turing (24.2 t/s) and Ada (46.0 t/s) put our 31-33 t/s exactly where our card class belongs - the decode deficit is the architecture, not our config. Their dual-GPU split also buys nothing, matching ours. Their 256K recipe (Q6_K + 4-bit KV) independently lands on the quantised-KV lever, and their pack confirms GGUF conversions carry no MTP head. Agentic: hard tier underway, first hard run passed.
08-15 09:30Third independent source added: a 254K context-depth sweep of 3.8-27B showing FP8 KV cache turns a -58% decode drop at depth into -2.4%. Our candidate loses ~31% decode at 131K on default KV, so a quantised-KV arm joins the extension list. It also shows the NVFP4 serve path working (speed only), softening but not lifting the no-NVFP4 flag. Agentic suite: 18 task-runs recorded, hard tier in progress.
08-15 08:57Tool calling: candidate 90.6 vs production 84.4 - first proven candidate win on a deciding axis (deterministic suite, 45/51 passed; misses cluster in trap scenarios, including one unconfirmed destructive delete). Card-bench complete on both configs; agentic suite now running.
08-15 08:20Second independent source added to the cross-check: a 34-entry community leaderboard has 3.8-27B as its overall quality #1, with the 3.6-35B keeping the decode crown - same split we measure. Their NVFP4 build of 3.8 is fully broken (0/69), so no NVFP4 arm here until that path matures.
08-15 08:05Every row now names a winner: filled tag = proven, outlined = current leader inside the noise band. Card-bench 2-GPU landed: candidate takes VRAM (17.6 vs 22.0 GiB at 8K), production takes speed (3.3-4.1x), power, and energy per token (~5.5x).
08-15 07:56Cross-check vs independent 10-task test added. Token-efficiency preliminaries: candidate leaner than production on the frozen suite (568 vs 674 tokens/item), with a cap-mismatch caveat now flagged for the repeat runs.
08-15 07:50Layout reworked to a single two-column head-to-head table (metrics as rows) for faster review.
08-15 07:45Page created. Quick-quality and speed-probe results for all four cells; card-bench running.
08-15 07:422-GPU quick-quality landed: composite 89.8. Candidate's own config spread (4.0 pts) now brackets production.
08-15 07:211-GPU quick-quality landed: composite 93.8, above both production cells.
08-15 06:58Chain started: triage gate passed, hypothesis registered, four phases queued on both GPU configs.