llm-lab bench campaign · 2x RTX PRO 4000 Blackwell
Last updated 2026-08-16 · SGLang Phase 1 final: NVFP4 2-GPU decode 57.32 t/s without speculation and 96.22 t/s with MTP (+68%), 1-GPU did not fit; the MTP arm is about 3x the llama.cpp Q4 serve but remains about 29% below production near 135 t/s, so Phase 2 stays off · addendum: Opus 4.6 now has quick-quality numbers via the same first-party claude -p path - corrected composite 94.4, the best on file (flags carried), knowledge 86.0 tops the table, needle 100 after a harness-cap artifact was diagnosed and remeasured · full Opus per-task tier detail remains its own section · CAMPAIGN CLOSED · Final verdict: keep production Qwen3.6-35B; Qwen3.8-27B lab-only (fails all three promotion clauses) yet the strongest quality candidate ever benched - matched agentic cell 42 vs 41 at the measured suite ceiling, best expert tier 16/18, tool-calling 90.6, maths +5.2, MTP +58% · Q8_0 quant arm final: Q4_K_M +56% decode, ~10.5 GB lighter · production restored and verified live (health 200) · eighteen sources on file
Production, defending
Qwen3.6-35B-A3B
Mixture-of-experts, 35B total / 3B active · Q4_K_M · the incumbent serving all live workloads
Candidate, challenging
Qwen3.8-27B
Dense 27B, thinking-native, coding + agentic positioning · Q4_K_M · released 2026-08-15
Quant-matched at Q4_K_M on identical hardware, both single-GPU and dual-GPU (layer-split) configurations. Same frozen eval sets, same serving stack.
One cell per axis, each with its final campaign read. Filled tag = proven separation. Outlined tag = measured leader whose gap remains inside the noise band (same-cell swing up to ~2.6 points on quality evals); the promotion verdict does not depend on those noise-band leads.
Every measured metric, one row each. Cells show both GPU configurations as 1 GPU / 2 GPU. Every row names its winner: a filled tag is a proven separation, an outlined tag is the current leader while the gap is still inside the noise band (winner judged on the mean of both configs, ~2.6-point same-cell swing on quality evals). Dead heat means the scores are literally identical. A purple-tinted Opus cell marks the two rows where Opus 4.6 beats both Qwens outright (composite and knowledge); on every other row at least one local model matches or beats the frontier reference.
| Metric (1 GPU / 2 GPU) | Qwen3.8-27B | Qwen3.6-35B | Read | Opus 4.6 (frontier ref) |
|---|---|---|---|---|
| Quick quality · frozen suite (matched 16K reasoning budget) | ||||
| Composite | 91.7 / 93.8 | 92.1 / 92.6 | Qwen3.8 +0.4 local · Opus best on file | 94.4 beats both Qwenscorrected · the run as printed said 69.4 because the needle category scored 0 on a harness-cap artifact (diagnosed and remeasured, see the needle row); corrected composite = (86.0 + 96.7 + 95.0 + 100.0) / 4 · best composite on file, flags below apply |
| Knowledge | 80.0 / 80.0 | 80.0 / 82.0 | Qwen3.6 +1.0 local · Opus 86 tops the table | 86.0 beats both Qwensbest on file · contamination-exposed for any frontier model |
| Maths | 97.0 / 100.0 | 93.3 / 93.3 | Qwen3.8 +5.2 · edges Opus by 1.8 on the mean | 96.70 errors · tools disabled, model-only arithmetic |
| Instruction-following | 90.0 / 95.0 | 95.0 / 95.0 | Qwen3.6 +2.5 local · Opus ties production at 95 | 95.0 |
| Long-context needle | 100.0 / 100.0 | 100.0 / 100.0 | Three-way 100 | 100.0remeasured · the first pass scored 0/6 at the suite's 64-token output cap because Claude Code treats its output cap as a hard ERROR, not a truncation point - and of the two answers that did return, one was harness identity confusion and one was Opus finding the buried needle and refusing to relay it as a suspected prompt injection; with the cap floored at 512 all six extractions came back exact |
| Refusal behavior · 14 calibrated probes per cell | ||||
| Refusal rate | 0.00 / 0.07 | 0.14 / 0.21 | Qwen3.8 wins local · Opus also 0.00 | 0.00 |
| Answered accuracy | 1.00 / 0.90 | 0.89 / 1.00 | Wash local · Opus 0.36 flagged | 0.3611 of 14 probes are accuracy-scored · Opus answered everything but agreed with the probe answer key far less often than the locals - with no temperature control and a different system prompt this is flagged, not concluded |
| Speed probe · 2K context, 3-run median after warmup | ||||
| Decode (t/s) | 31.4 / 32.8 | 139.7 / 142.0 | Qwen3.6 4.4x | n/a · cloud |
| Prefill (t/s) | 844 / 873 | 1057 / 1048 | Qwen3.6 1.2x | n/a · cloud |
| TTFT p50 (ms) | 611 / 591 | 488 / 493 | Qwen3.6 wins | n/a · cloud |
| SGLang Phase 1 · NVFP4 weights, bench port, 2-GPU tensor parallel · Phase 2 off | ||||
| Decode, no-spec / MTP (t/s) | 1 GPU: no fit · 2 GPU: 57.32 / 96.22MTP +68% over the SGLang no-spec arm and about 3x the candidate's llama.cpp Q4 serve | ~135production Qwen3.6-35B MoE | Qwen3.6 ~1.4x vs MTP | n/a · cloud |
| Prefill, no-spec (t/s) | 1 GPU: no fit · 2 GPU: 22,283.5 | not matched in this Phase 1 probe | SGLang wins candidate prefill | n/a · cloud |
| Card-bench · 50%-of-context prompts, 3-run median · both configs in | ||||
| Decode at 8K context (t/s) | 30.6 / 31.9 | 128.4 / 130.1 | Qwen3.6 4.1x | n/a · cloud |
| Decode at 131K context (t/s) | 21.2 / 22.6 | n/a / 77.0 | Qwen3.6 3.4x | n/a · cloud |
| Concurrent 4-stream aggregate (t/s) | 68.9 / 77.1 | 243.8 / 252.4 | Qwen3.6 3.4x | n/a · cloud |
| VRAM peak at 8K (MiB) | 16,834 / 17,616 | 21,504 / 22,034 | Qwen3.8 wins | n/a · cloud |
| Power draw at 8K (W) | 150.4 / 224.8 | 127.8 / 165.6 | Qwen3.6 wins | n/a · cloud |
| Energy per token at 8K (J/tok) | 4.9 / 7.1 | 1.0 / 1.3 | Qwen3.6 ~5x | n/a · cloud |
| KV-cache precision A/B · candidate weights, same 3-run methodology · lab default is q8_0 KV (a this-campaign discovery: every published card-bench number is a q8-KV number) | ||||
| Decode at 131K by KV type (t/s, 1 / 2 GPU) | f16: crash / 24.9 · q8: 21.2 / 22.6 · q4: 20.9 / 22.3 | n/a · diagnostic | f16 +10% where it fits | n/a |
| Prefill at 131K by KV type (t/s, 2 GPU) | f16: 49.7 · q8: 41.2 · q4: 42.3 | n/a · diagnostic | f16 +21% | n/a |
| VRAM peak at 131K by KV type (MiB, 2 GPU) | f16: 26,496 · q8: 23,596 · q4: 21,548 | n/a · diagnostic | q4 wins, -4.8 GB vs f16 | n/a |
| 131K on a single 24 GB card | f16 KV does not fit (server crash) · q8: 21.5 GiB ok · q4: 19.0 GiB ok | n/a · diagnostic | quantised KV required | n/a |
| Weight-quant A/B · Q4_K_M vs Q8_0 on the same weights · 2-GPU only (Q8_0 is 27.1 GiB and cannot fit one card), same 3-run methodology | ||||
| Decode at 8K by weight quant (t/s, 2 GPU) | Q4_K_M: 31.9 · Q8_0: 20.4 | n/a · diagnostic | Q4 +56% | n/a |
| Decode at 131K by weight quant (t/s, 2 GPU) | Q4_K_M: 22.6 · Q8_0: 16.2 | n/a · diagnostic | Q4 +40% | n/a |
| Prefill at 131K by weight quant (t/s, 2 GPU) | Q4_K_M: 41.2 · Q8_0: 37.7 | n/a · diagnostic | Q4 +9% | n/a |
| VRAM peak by weight quant (MiB, 2 GPU, 8K / 131K) | Q4_K_M: 17,616 / 23,596 · Q8_0: 28,396 / 34,376 | n/a · diagnostic | Q4 saves ~10.5 GB | n/a |
| Tool calling · 51 deterministic scenarios, temperature 0 | ||||
| Tool-calling /100 | 90.6 | 84.4 | Qwen3.8 +6.2 | not run |
| Agentic · complete (64K context, matched to production's reference) | ||||
| Agentic suite /100 (weighted) | 71.1 | 83.5 | Qwen3.6 +12.4 | n/a · harness-specific |
| Agentic 64K context /45 (shared 4096-token output budget) | 35 | 41 | Qwen3.6 wins | n/a · no cap applies |
| Agentic /45 at raised 16K output cap (diagnostic, reported separately) | 39expert 15/18 (was 11/18) · hard 9/12 (hard_base 0/3 by clean 630s timeouts) · medium/small/trivial 15/15 | 42on file, same cap | Qwen3.6 wins | n/a · no cap applies |
| Agentic /45 with thinking OFF at the matrix 4096 cap (matched cell) | 42FINAL · expert 16/18 (thinking-on matrix: 11/18) · hard 11/12 · medium/small/trivial 15/15 · zero truncations, zero loops | 41measured 08-15 19:22 in the identical config (thinking off, 4096 cap, 64K ctx) · expert 14/18: expert_calc 1/3, expert_money 1/3 · hard 12/12 · medium/small/trivial 15/15 · production's best anywhere on file stays 42 (thinking on, 16K cap) | Qwen3.8 wins the matched cell | ≈41same 19 tasks, hidden oracles, via first-party claude -p · full 3-run cells now on every expert and hard task: expert 14/18 (expert_calc 1/3, expert_money 1/3 - production's exact failure profile; the other four experts all 3/3), hard 12/12, all 24 tier-repeat runs passed at 12-34 s wall · with single-run medium/small/trivial 15/15 that lands ≈41/45 - identical to production's split, one under the candidate's 42 |
| Agentic /45 with thinking capped at 2048 (diagnostic) | 39FINAL · expert 15/18 (expert_money 2/3 where thinking-off went 3/3) · hard 9/12: hard_base 0/3, all three 630s wall-clock timeouts mid-work, no loops - per-turn thinking starves the time budget on the longest-horizon task (thinking-off went 2/3) · medium/small/trivial 15/15 | 42same comparators as the row above | Qwen3.6 wins | n/a · no cap applies |
| Token efficiency · quick-quality suite, output tokens per scored item (matched 16K reasoning budget) | ||||
| Tokens per item (all scored) | 765 / 906 | 674 / 725 | Qwen3.6 ~16% leaner | n/a · not comparable |
| Tokens per correct answer | earlier 633 vs 768 candidate lead withdrawn - it compared a capped run against an uncapped one | superseded by the matched all-scored row above | Cap artifact | n/a |
| Tokens-to-answer, agentic tasks | not run - the matched-cap quick-quality row above settled the token-efficiency axis (production ~16% leaner) | not run | Closed unrun | n/a |
The matched-budget quality repeats are complete: the candidate leads the local composite by +0.4 on the mean, inside the measured ~2.6-point noise band, while production is ~16% leaner per scored item. Card-bench gives the candidate a clean footprint win - 4.3-4.7 GiB less VRAM at 8K on both configs - but production keeps 3.4-4.1x speed and roughly 5x better energy per token, and every candidate cell ran throttled against the software power cap. On tool calling the candidate's +6.2 comes from a deterministic temperature-0 suite (45 of 51 scenarios passed), so it is tagged as a win, not noise; its six misses cluster in trap scenarios, including one that matters operationally - it fired a destructive delete with confirm=true without asking the user first. The speed gap is the expected dense-vs-MoE cost: minor for chat (32 t/s outruns reading speed), compounding for thinking chains and agentic loops. The Opus 4.6 column is a frontier reference, not a contender: it ran the same 19 tasks against the same hidden oracles via first-party claude -p on cloud hardware, so pass/fail transfers but speed, VRAM, power, and token metrics do not. Quick-quality has now been run for Opus through the same first-party path (a driver that swaps the suite's chat call for a claude -p subprocess), and its numbers carry their own flags: the Claude Code harness system prompt is injected, there is no temperature or seed control (locals ran temp 0, seed 26), tools are disabled, knowledge and maths are contamination-exposed for any frontier model, and the needle category's output cap had to be floored at 512 tokens against the locals' 64 - output-budget parity is impossible below 512 on this path. Its agentic result is the calibration that reframes the whole /45 axis: 17/19 on the first pass, failing only the exact two expert tasks production drops, with repeats pinning both at 1/3 - so the suite's effective ceiling is about 41-42/45 and both local finalists are already sitting on it.
The frontier reference, opened up. Every expert and hard task carries a full 3-run Opus cell (24/24 tier-repeat runs completed; the two failure cells were repeated in the earlier calibration pass). Comparators are the matched thinking-off cells for both local models. Medium, small, and trivial tiers were 15/15 for all three models and are collapsed into one row; Opus ran those once, not three times. Wall-clock is not comparable across the columns - Opus ran on cloud hardware through a first-party CLI harness (repeat runs landed in 12-34 seconds per task) - so this table reports pass counts only.
| Task | Qwen3.8-27B (think-off) | Qwen3.6-35B prod (think-off) | Read | Opus 4.6 |
|---|---|---|---|---|
| Expert tier · 6 tasks × 3 runs | ||||
| expert_calc | 1/3 | 1/3 | Nobody clears it | 1/3 |
| expert_money | 3/3 | 1/3 | Only Qwen3.8 clears it | 1/3 |
| expert_csv_dialect | 3/3 | 3/3 | Three-way clean | 3/3 |
| expert_debug_eventbus | 3/3 | 3/3 | Three-way clean | 3/3 |
| expert_reservations | 3/3 | 3/3 | Three-way clean | 3/3 |
| expert_scheduler | 3/3 | 3/3 | Three-way clean | 3/3 |
| Expert tier /18 | 16 | 14 | Best expert tier on file | 14 |
| Hard tier · 4 tasks × 3 runs | ||||
| hard_base | 2/3 | 3/3 | Qwen3.8 gives one back | 3/3 |
| hard_duration | 3/3 | 3/3 | Three-way clean | 3/3 |
| hard_intervals | 3/3 | 3/3 | Three-way clean | 3/3 |
| hard_lru | 3/3 | 3/3 | Three-way clean | 3/3 |
| Hard tier /12 | 11 | 12 | Opus matches production | 12 |
| Lower tiers · medium + small + trivial | ||||
| Medium / small / trivial /15 | 15 | 15 | Saturated for all three | 15single-run for Opus |
| Total /45 | 42 | 41 | Qwen3.8 by one, inside noise | ≈41 |
The reportable headline: a frontier model, run over the same tasks, sandbox, and hidden oracles, produces production's tier profile exactly - down to failing the same two expert tasks at the same 1/3 rate - and lands one point under the local candidate's matched cell. The two structural cells are expert_calc, which no model clears at better than 1/3 in any configuration (task-hard, effectively excluded from the reachable ceiling), and expert_money, where Qwen3.8 thinking-off is the only model of the three to score above 1/3. That is why the suite's effective ceiling is 41-42/45 and why the /45 axis stops discriminating at the top. What this table cannot say: anything about speed, cost, or tokens (different harness and hardware). Quick-quality has since been run for Opus over the same first-party path - those numbers, with their flags, live in the head-to-head table above.
Eighteen external sources so far. Source 1: a published 10-task head-to-head of Qwen 3.8 27B vs Qwen 3.6 27B (dual RTX 6000 Pro, BF16, multi-token prediction enabled). Source 2: a live 34-entry community leaderboard running its own task suite across 3090 and DGX Spark clusters. Source 3: a context-depth decode sweep of Qwen 3.8 27B NVFP4 on 4x RTX 3090, comparing FP8 vs BF16 KV cache out to 254K. Source 4: llama-bench decode numbers for the same GGUF quants on Turing and Ada workstation cards - the closest stack to ours. Source 5: a ts-bench coding-agent run of the UD_Q4_K_XL GGUF with thinking off, on llama.cpp like ours, plus the same author's follow-up forensic trace of its loop failures. Sources 6 and 7: serving recipes from one author (DGX Spark + RTX 6000 PRO on vLLM, then a single RTX 5090 on SGLang), configuration detail but no benchmarks - the 5090 repo has since added a work-in-progress banner and a worked 32GB memory budget, still without a single measured number. Source 8: a single-DGX-Spark deployment under a vLLM 0.27 nightly, and the first repo to publish hard multi-token-prediction numbers for this model. Source 9: the vendor's own published benchmark table (via a widely shared post) - marketing numbers at full precision, held to a lower evidentiary standard than the independent sources. Source 10: the official vLLM serving recipe for the model - authoritative configuration, no benchmarks. Source 11: a measured MTP speculative-decoding A/B on llama.cpp with the same unsloth Q4_K_M GGUF we run, on 24GB cards. Source 12: a published CRUDbench comparison (X, @LottoLabs, 2026-08-15) of the model's non-thinking vs highest-effort thinking modes on a tightly specified deletion task. Source 13: a community NVFP4 quantization with the MTP draft head preserved in bf16, with measured vLLM throughput across concurrency levels on four 16GB Blackwell cards. Source 14: relayed vendor deployment guidance on reasoning-effort levels and long-context extension. Source 15: an Intel Arc Pro B70 inference cookbook with a measured MTP depth ladder (vLLM XPU, GPTQ-Int4 weights, BF16 MTP head) - self-reported, held provisional. Source 16: a thinking-budget sweep on an HTML canvas generation test (X, 2026-08-16) - qualitative video comparison, no scores. Source 17: a practitioner report preferring a medium thinking budget standalone and no-thinking execution under an external planner (X, @TheAhmadOsman, 2026-08-16) - anecdote. Source 18: the official SGLang Qwen3.8-27B deployment cookbook, including the exact H200 FP8 balanced cell and architecture-specific guidance - authoritative configuration, no published speed table. Different hardware and precision: ratios and direction should transfer even where absolutes cannot.
| Their finding | Our measurement | Verdict |
|---|---|---|
| Source 1 · 10-task head-to-head, dual RTX 6000 Pro, BF16 + MTP | ||
| 3.8 wins 9/10 real tasks vs 3.6-27B | Direction agrees at matched budget on quick quality (91.7/93.8 vs 92.1/92.6, candidate +0.4 mean) and the agentic arbiter is now in: thinking-off the candidate takes the matched cell 42 to 41 | Consistent |
| 3.8 decodes 1.9x faster than 3.6-27B (93 vs 50 t/s) | Plain llama.cpp Q4 is effectively tied with its 3.6-27B sibling on our stack, but the mechanism transfers: our llama.cpp MTP arm reaches 49.8 t/s (+58%), and SGLang NVFP4 MTP reaches 96.22 t/s on 2 GPUs. Against our faster 35B MoE production model near 135 t/s, the candidate still trails by about 29% | Mechanism reproduced |
| 3.8 uses 2.9x more tokens per task | Direction now reproduces at matched budget, at far smaller magnitude: 3.8 spends ~16% more tokens per scored item than production (766/906 vs 674/725). The earlier "3.8 leaner" read was an artifact of the candidate's tighter output cap. His 2.9x came from long agentic jobs, where the thinking channel compounds the appetite | Reproduced in direction |
| Source 2 · 34-entry live leaderboard, 3090 / DGX Spark clusters, vLLM + llama.cpp | ||
| 3.8-27B (INT4, 2x 3090) is their overall #1 on quality (97.1), above every Qwen3.6 variant, MiMo-V2.5, GLM-5.2 and DeepSeek-V4 Flash | Same direction as our quality result, stronger magnitude. Their metric weights quality 70/30 over speed, so it structurally favors 3.8. Our completed agentic /100 comparison did not support promotion | Quality direction consistent |
| Qwen3.6-35B-A3B still owns raw decode on their boards (up to 178 t/s) | Matches our proven speed split: production 3.3-4.4x faster on every throughput axis | Consistent |
| Their NVFP4 build of 3.8 scores 0.0 with 69/69 task failures | We run GGUF Q4_K_M, unaffected. Flag: the NVFP4 quant path for this model looks broken right now, so no NVFP4 arm until that matures | Noted, avoided |
| Source 3 · 72-point context-depth sweep, 4x RTX 3090, NVFP4 weights, FP8 vs BF16 KV cache | ||
| FP8 KV cache holds decode nearly flat at depth: -2.4% at 254K vs -58% with BF16 KV | A/B CLOSED on our stack, and the direction is the opposite of their claim. Card-bench's default was already q8_0 KV, so the 30.6-to-21.2 depth slide happens WITH quantised KV. The chained true-f16 baseline then measured FASTER at depth, not slower: f16 KV decodes 24.9 t/s at 131K on 2 GPUs vs q8's 22.6 (+10%) and prefills 49.7 vs 41.2 (+21%) - on our llama.cpp path, quantising KV costs depth speed, it does not recover it. What quantised KV buys is fit: f16 KV cannot hold 131K on one 24 GB card (server crash, 25.9 GiB peak on 2 GPUs), and q4_0 saves a further ~2 GB over q8 at no meaningful speed cost. Most of the depth slide is attention compute either way - even f16 falls 32.1 to 24.9 from 8K to 131K | Reversed on our stack |
| 3.8-27B NVFP4 serves fine via a Marlin kernel, even on GPUs with no native FP4 or FP8 | Counter-evidence to Source 2's broken NVFP4 entry, now reinforced by our own clean SGLang Phase 1 serve. Both are speed evidence only, not quality evidence, so the NVFP4 checkpoint still cannot support a promotion claim | Serve path reproduced |
| Source 4 · llama-bench tg128, same GGUF quants, Turing 24GB (SM75) vs Ada 48GB (SM89) | ||
| Q4_K_M decodes 24.2 t/s on a 2018 Turing card and 46.0 on Ada | Our 31-33 t/s sits exactly where our card class belongs on that ladder, on the same engine and quant. Confirms the 4x deficit vs production is the dense-27B architecture, not a serve misconfiguration on our side | Consistent |
| Dual-GPU layer split buys nothing: 46.5 vs 46.0 t/s single-card | Matches our split exactly (31.9 vs 30.6 across configs, within noise). Third stack to show single-stream decode does not scale with a second card | Consistent |
| Q8_0 (26.6 GiB) does not load on a 24GB card; their 256K recipe is Q6_K plus 4-bit KV cache | The fit prediction transferred: our 27.1 GiB Q8_0 arm was 2-GPU only. It also settled the tradeoff in Q4_K_M's favour: Q4 decoded 40-56% faster and used about 10.5 GB less VRAM. Quantised KV remained the fit lever at 131K, where true f16 KV could not fit one 24 GB card | Fit prediction reproduced |
| Source 5 · ts-bench coding-agent run, UD_Q4_K_XL GGUF, thinking off, llama.cpp | ||
| 3.8-27B is significantly worse as a coding agent than 3.5 or 3.6, getting stuck in infinite loops until automatic compaction | Same direction as our decided agentic axis (35/45 vs production's 41/45), and the same signature: our harness tagged every failed stall "looped" before the trace showed they are 4096-token cap truncations. A second llama.cpp stack seeing agent loops that 3.5/3.6 never produced makes the weakness look real, not local | Consistent |
| Their loops happen with thinking OFF | Not reproduced on our stack. The completed thinking-off arm scored 42/45 with zero truncations and zero loops, versus 35/45 with thinking on at the matched cap. Their long-context loop hazard remains useful external evidence, but it did not predict our best serving mode | Not reproduced locally |
| They initially suspected a chat-template error or llama.cpp optimisation issue rather than the model | Their own follow-up (rows below) now points away from serving: tool outputs were verified correct with zero API errors across all three looping sessions. On our stack too, the embedded template passes Phase-0 probes and scored 90.6 on tool calling | Superseded below |
| Follow-up forensics: the loops are exact-command repetition - one task re-ran the identical string-concat command 908 times over 2h43m, receiving a correct result every time, and still could not update its judgment (another: 500 identical runs over 1h45m) | A different member of the same failure family as our stalls: ours died in a single turn at the output cap, theirs died across hundreds of turns of correct-but-ignored tool results. The completed raised-cap and thinking-off arms showed no repetition signature on our workload, but the external long-trajectory hazard remains operationally relevant | Not reproduced locally |
| After automatic compaction the same tasks pass immediately (14/14 in 23 seconds) - long accumulated context breaks it, fresh context works; the compaction summary even misattributes the failure to "broken" tool results that were in fact correct | Strong model-side signal: in-context judgment degrades as the transcript grows, and the model cannot see its own degradation. Our matrix runs sit at 64K matched context; production 3.6 has never shown this mode on the same suite. This goes in the defended verdict as an operational-hazard finding, independent of the cap artifact | Model-side signal |
| Source 6 · MiaAI-Lab serving recipe, DGX Spark GB10 (128GB unified) + RTX 6000 PRO (96GB), vLLM NVFP4 | ||
| FP8 KV cache doubles KV capacity on their rig: 1.14M to 2.30M cached tokens, 2x concurrency at 1M context | Our completed 131K A/B confirms the capacity direction: q8 KV fits one 24 GB card where f16 crashes, and q4 saves about another 2 GB. The speed direction did not transfer - true f16 was faster where it fit - so quantised KV is a capacity lever here, not a decode-speed lever | Capacity direction reproduced |
| Working NVFP4 + FP8-KV + MTP (2 speculative tokens) recipe on vLLM, 1M context via 4x YaRN | Third data point that the NVFP4 serve path is maturing (after Source 2's broken 0/69 build and Source 3's Marlin workaround), and the first showing MTP running for this model in an engine. But it is a 96-128GB memory class; the 1M-context recipe cannot fit our 24GB cards, and they publish no throughput or quality numbers at all | Recipe only |
| No benchmarks: the repo ships configuration and capacity figures, not speed, quality, or agentic scores | Its value was operational: it showed that the MTP head survives in the NVFP4 checkpoint. Our SGLang Phase 1 run has now independently verified that path on our hardware and supplied the missing speed numbers; the source itself still contributes no benchmark result | Path independently verified |
| Source 7 · MiaAI-Lab serving recipe, RTX 5090 32GB (SM120 Blackwell), SGLang NVFP4 | ||
| MTP speculative decoding on SM120 Blackwell: SGLang + FlashInfer using the checkpoint's built-in MTP head | Phase 1 is now measured on our cards. The 1-GPU arm did not fit. On 2 GPUs, no-spec decoded at 57.32 t/s and MTP at 96.22 t/s, a +68% lift; no-spec prefill reached 22,283.5 t/s. The best arm is about 3x the candidate's llama.cpp Q4 serve, but it remains about 29% below production near 135 t/s | Measured: Qwen3.6 still wins speed |
| Reasoning depth is tunable per request (xhigh, medium, low - xhigh is the default) with thinking preserved across turns | The completed dose-response settles this on our workload: thinking-on scored 35/45 at the matrix cap, raised-cap and capped-thinking each scored 39/45, and thinking-off scored 42/45. The deepest default explains the token appetite; effort-down helped, but off was best | Dose response measured: off wins |
| Hybrid-GDN serving quirks documented: 2048-token prefill chunks (8192 causes ~600ms decode stalls), 78.4MB state per concurrent slot, no benchmarks published | Recipe only, no verdict input. The per-slot state cost explains why this architecture concurrency-limits hard on small VRAM, and the chunking trap is worth carrying into any future engine trial on our side | Noted |
| Repo update (08-16): the headline now reads work in progress - does not work currently, and it adds a worked 32GB memory budget: 17-22GB weights, 78.4MB hybrid state per slot with 8 slots per request (4 lazy + 4 MTP draft), concurrency 2 via a 1.25GB mamba-cache pool, fp8 KV at 32.8KB/token, context capped at 100K, prefill chunked at 2048 (8192 chunks stall decode ~600ms on this architecture) | Its implementation status remains weak external evidence, but its fit warning transferred: our 1-GPU arm did not fit. The 2-GPU Phase 1 arm did boot and measure cleanly, so the external recipe is no longer carrying the feasibility claim by itself. Phase 2 remains off because the measured MTP arm still trails production decode by about 29% | Fit warning reproduced |
| Source 8 · Single DGX Spark (GB10), vLLM 0.27 nightly, NVFP4 + FP8 KV + MTP depth 3 | ||
| First hard MTP numbers for this model: speculative decoding at depth 3 takes single-stream decode from 11.1 to 31.7 t/s (2.85x); depth 4+ crashes because the checkpoint carries only one MTP layer, and ngram speculation is ineffective for this workload | The direction transferred but the ratio did not: our SGLang Phase 1 arm moved from 57.32 to 96.22 t/s, a 1.68x lift. That is enough to make SGLang about 3x faster than the candidate's llama.cpp Q4 serve, but not enough to catch production near 135 t/s. The source correctly predicted a useful MTP gain; our measurement closes the promotion question in production's favour | Direction reproduced, smaller gain |
| NVFP4 beats FP8 on both throughput axes on their rig: 31.7 vs 21.9 t/s single-stream, 313 vs 225 t/s aggregate at 64-way concurrency, via a native FlashInfer kernel | Our SGLang Phase 1 result confirms that NVFP4 can serve quickly on this hardware class, but it remains a speed-only result. No source has yet shown this NVFP4 checkpoint scoring cleanly on quality, so it cannot replace the quant-matched Q4 evidence in the promotion decision | Speed path reproduced |
| Tool calling requires the qwen3_xml parser; the JSON-based hermes parser fails silently. First large prefill stalls ~13s on kernel JIT until warmed | Our stack is unaffected - llama.cpp's embedded template with --jinja passes the Phase-0 probes and scored 90.6 on the tool suite - but a silently failing default parser is a plausible root cause for some community "3.8 cannot tool-call" reports, and both traps go in the notebook for any vLLM or SGLang trial on our side | Noted |
| Source 9 · Vendor benchmark table, full precision (marketing numbers, lower evidentiary standard) | ||
| Headline axis is agentic coding: Terminal Bench 73.0 vs 63.4, SWE-bench Pro 61.7 vs 53.5, QwenSWEBench 79.0 vs 49.3 against the 3.6-27B sibling, beating a frontier lab model on several rows | Their comparator is the dense 3.6-27B, not our production 3.6-35B MoE, and they run full precision against our Q4_K_M - so the bars differ. On our suite the story is conditional: thinking-on the candidate loses to production at both caps (35 and 39 vs 41 and 42), thinking-off it wins the matched cell 42 to 41. The vendor's "large agentic upgrade" survives contact with our data only in the thinking-off configuration | Different bar |
| General-knowledge gains are modest in their own table: GPQA 89.2 vs 87.8 for the 27B sibling | Matches our knowledge axis, where production keeps a narrow lead at matched budget (+1.0 mean: 82 vs 80 two-card, 80-80 tie one-card; the old +4.0 included a budget-starved 74): even the vendor does not claim general knowledge as the upgrade, which is consistent with agentic-and-coding being where this checkpoint spent its training budget | Consistent |
| Tool-use and instruction rows are strong: IFBench 79.5 vs 69.1, LiveCodeBench 90.3 vs 83.9 | Direction reproduces on our deterministic tool suite (90.6 vs production's 84.4, the candidate's one proven deciding-axis win). The instruction-following claim we have not measured separately; it rides inside the agentic and quick-quality scores | Partly reproduced |
| Source 10 · Official vLLM serving recipe (authoritative configuration, no benchmarks) | ||
| NVFP4 weights occupy 24.6 GiB on a single Blackwell GPU at tensor-parallel 1 | Hard gate input for our speculative-decoding probe: 24.6 GiB of weights alone exceeds one of our 24GB cards before any KV cache, so a single-card NVFP4 serve is out - the probe would need both cards, which costs the production-parallel angle. The VRAM-fit gate the probe carries is now answered in the negative for single-card | Probe gate tightened |
| Official speculative-decoding config ships MTP at exactly 3 speculative tokens | Converges with Source 8's measurement from the other direction: they found depth 3 optimal empirically (2.85x) and depth 4+ crashing on the single MTP layer; the vendor recipe simply never offers more than 3. The probe config is settled | Converges |
| Thinking controls are first-class serving options: enable_thinking false, and reasoning_effort low / medium / xhigh via chat-template kwargs | Direct support for our final serving recommendation: thinking-off scored 42/45 and the completed capped-thinking arm did not beat it. The official SGLang cookbook now independently names qwen3_coder as the parser for this checkpoint, resolving the earlier qwen3_xml version-drift caveat for that engine | Supports our measured best config |
| Source 11 · llama.cpp draft-MTP A/B, same unsloth Q4_K_M GGUF, RTX 3090 24GB + RTX 5090 mobile 24GB | ||
| MTP speculative decoding works in llama.cpp itself (the July draft-mtp support) with the quantizer-preserved draft head, measuring +33% decode on a 3090 (31.0 to 41.3 t/s, acceptance 0.76-0.82) and +39% on a 5090 mobile | Verified both prerequisites live on our stack within the hour: our bench binary already exposes the draft-mtp option, and our exact GGUF carries the four draft-head tensors. Their 3090 baseline of 31.0 t/s is nearly identical to our 31-33, so their +33% is a direct prediction for our cards - no engine change, no requantization, no second GPU. A same-binary MTP A/B jumped the queue - and it is now MEASURED on our cards: baseline 31.5 t/s, draft-MTP n-max 2 = 49.8 t/s, a +58% decode lift at 71.8% acceptance (3-run medians, warmup discarded, greedy 512-token completion, single slot). That beats their +33% prediction on a near-identical baseline. The May caution resolved exactly as their numbers claimed: Qwen3.6's weak draft head (0.67 acceptance, 12% slower) was the model's fault, not the engine's. Also corrects Source 4's "GGUF packs carry no MTP head": that is pack-dependent, and the unsloth pack we downloaded kept it | Confirmed on our stack: +58% |
| Draft-depth tuning is shallow on their cards: n-max 2 wins (50.9 t/s), n-max 3 drops to 48.3 with acceptance falling 0.76 to 0.68, prose degrading faster than code | Engine-dependent optimum: the vLLM and SGLang sources converge on depth 3, llama.cpp's sweet spot is 2. Our A/B ran both and REPLICATES the ordering: n-max 2 = 49.8 t/s at 71.8% acceptance, n-max 3 = 48.8 t/s at 64.7% acceptance - depth 3 drafts more tokens (521 vs 419) but accepts a lower fraction, exactly their pattern. Their methodology also mirrors ours (3-run medians, warmup discard) and was measured at 131K resident context on quantised KV - yet another co-occurrence of the quantised-KV lever our extension arm tests | Replicated: n-max 2 wins |
| Source 12 · CRUDbench contract task, non-thinking vs xhigh thinking (X post, @LottoLabs, 2026-08-15) | ||
| On a tightly specified find-check-delete task, non-thinking stayed literal to the contract and passed all three hidden fixtures in 27.7s with 5 model calls, 931 output tokens, and zero reasoning characters; xhigh inferred soft-deletion and status semantics from schema fields nobody asked about, failed, and retries kept reintroducing the invented conditions | Same direction as our ladder (thinking-off 42/45 vs thinking-on 35/45) and it supplies the mechanism our truncation trace could not see: the reasoning channel does not just overrun budgets, it invents requirements. The retry behavior - correcting one assumption then reinstating it - also echoes Source 5's judgment-collapse signature, where correct evidence fails to update the model's beliefs. The author's framing is careful and matches ours: not "thinking bad", but reasoning depth changes operating behavior, and extra inference only helps when the inferences are correct | Consistent, adds mechanism |
| Their conclusion: for tightly specified tasks, non-thinking better enforces the stated contract | Now a three-source convergence: their contract test, our 19-task suite, and the vendor shipping thinking-off as a first-class mode. The completed capped-thinking arm scored 39/45 against thinking-off at 42/45, so the evidence supports thinking-off as this deployment's default posture | Convergent: off wins |
| Source 13 · community NVFP4 quantization with the MTP head preserved (vLLM, four 16GB Blackwell cards, TP4, 128K ctx, fp8 KV) | ||
| MTP at 3 speculative tokens lifts single-stream decode 49.0 to 72.6 aggregate t/s (+48%), and the gain survives concurrency: still +22% at 8-way (318.3 to 386.9) | The single-stream direction reproduced twice on our stack: llama.cpp gained +58% at n-max 2, while SGLang gained +68% with EAGLE 3/1/4. The optimum remained engine-specific - llama.cpp n-max 3 slipped below n-max 2 as acceptance fell. We did not run a matched concurrency arm, so their positive 8-way result remains external evidence rather than a local claim | Single-stream gain reproduced |
| Their NVFP4 pack is 20.6 GB (bf16 55.6 GB compressed W4A4, with the MTP head, vision tower, and output layer kept bf16), against the official recipe's 24.6 GiB footprint | Partially reopens the single-card question Source 10 closed: 20.6 GB of weights fits one of our 24GB cards, though the ~3 GB left for KV and activations is thin at any useful context. Caveat on their platform: W4A16 does not serve at all on this architecture in their vLLM version (kernel tile mismatch), so W4A4 is the only NVFP4 flavor that runs - quality of W4A4 group-16 on this model is unverified by anyone so far | Reopens single-card, quality unverified |
| Their gotchas: the bf16 draft head must stay on the quantization ignore-list or acceptance silently drops to 0% and MTP goes net-negative; the model needs a generation budget of at least 4096 or the thinking phase eats it; tool parser qwen3_xml on their vLLM version | All three converge with our findings: the silent-degradation failure mode is the exact trap our A/B methodology (acceptance-rate logging, 3-run medians) is built to catch; the budget warning is our cap-starvation result restated by an independent operator; and the parser naming settles the Source 8 vs Source 10 drift - qwen3_xml on older vLLM, qwen3_coder on the current recipe, version-dependent not contradictory | Convergent on all three |
| Source 14 · vendor deployment guidance, reasoning effort and long context (relayed) | ||
| Vendor warning: in multi-turn agentic tasks, LOWER reasoning effort can raise total latency and tokens through more failures and retries (xhigh is the default effort) | Our completed arm measured the opposite on this workload: capped thinking scored 39/45 and thinking-off scored 42/45 with zero truncations and zero loops, versus 35/45 with thinking on. The disagreement remains workload- and precision-scoped rather than universal, but our data governs this deployment | Our measured arm favors off |
| Native context 262K; YaRN static scaling extends to 1M, with the factor tuned to typical prompt length rather than maximum | Matches Source 10's recipe (262K native, ~1M via overrides) - consistent, nothing new for us since our suite runs at 64K and the extension arm targets 131K resident, both inside native range with no YaRN needed. The tune-to-typical-length advice is worth keeping if a long-context serve ever happens | Consistent, not load-bearing |
| Source 15 · Intel Arc Pro B70 inference cookbook (vLLM XPU, GPTQ-Int4 weights + BF16 MTP head, 230W card) - self-reported, provisional | ||
| MTP depth ladder at p512/g128 decode: no speculation 32.9 t/s, 1 draft token 52.0 (+58%), 2 tokens 65.8 (+100%), 4 tokens 83.7 (+154%), acceptance 93.7-100% | Their +58% at one draft token is numerically identical to our llama.cpp n-max 2 measurement, from a different engine, quant format, and silicon vendor - a striking replication of the first rung. But their ladder keeps paying at deeper drafts where ours reversed (our n-max 3 lost ground as acceptance fell to 64.7%), so the depth optimum is engine-dependent, not universal: llama.cpp saturates at 2, vLLM-family engines keep gaining to 3-4 while acceptance stays above ~90%. Refines Source 11's "2 beats 3" from a model fact to an engine fact | +58% replicated, depth optimum is engine-dependent |
| Full-context run (130,944-token prefill): decode 23.2 t/s without speculation, 56.3 with 4 draft tokens (+143% at depth) | First source to show the MTP gain surviving depth. Their un-specced 23.2 t/s at full context is almost exactly our 131K figure (22.6), and they recover 2.4x of it - if that transferred to any engine our cards run, the candidate's worst axis at depth would close substantially. Single self-reported run on Intel silicon, so it stays provisional until reproduced | Provisional, high value if it transfers |
| Source 16 · Thinking-budget sweep on an HTML canvas test (X, @KyleHessling1, 2026-08-16) - video comparison, qualitative labels, no scores | ||
| Reasoning-budget ladder for Qwen 3.8 27B: thinking OFF and 2K caps labelled broken, 6K best, 12K excellent, 24K wasteful; wall time climbs ~150s to ~550s across the sweep; claimed sweet spot 6-12K; notes the model defaults to maximum thinking | Half confirms, half inverts our data - and the split is the finding. The "defaults to maximum thinking, responds well to caps" claim matches our dose-response ladder exactly (thinking length is a controllable cost lever on this model). But the quality ordering is the inverse of ours: our matched agentic cell peaked at thinking OFF (42/45) and every thinking dose cost 3-7 points, while his OFF is broken and quality peaks mid-budget. Reconciliation is workload, not contradiction: single-shot codegen where any mistake breaks the render rewards CoT, agentic tool-loops punish it. So the serve recommendation becomes conditional - thinking off for agentic serving (our measurement), a 6-12K cap rather than unbounded for single-shot codegen (his). Single-task, no repeats, no scores | Domain-split: confirms the cap lever, inverts the OFF verdict |
| Source 17 · Practitioner report on thinking budget and role-splitting (X, @TheAhmadOsman, 2026-08-16) - anecdote, no numbers | ||
| Claims Qwen 3.8 27B "improves tremendously" at a medium thinking budget, and that his preferred setup is another model doing the thinking and planning while 27B runs implementation with no thinking at all | The second half is our agentic finding restated as a user preference: when planning lives outside the model - a planner model, or an agent harness driving a tool loop - the executor runs best thinking-off, which is exactly the cell where the candidate scored its 42/45. The first half lands beside Source 16's mid-budget sweet spot for standalone use. Three independent lines of evidence (our measured ladder, Source 16's sweep, this report) now triangulate one picture: thinking budget is a real quality lever on this model, the default is too high, and OFF is correct when the model is the executor rather than the planner | Anecdote: corroborates the workload split |
| Source 18 · Official SGLang Qwen3.8-27B cookbook, H200 FP8 balanced cell and Blackwell NVFP4 guidance | ||
| The H200 FP8 balanced cell is a verified single-node deployment recipe: FP8 weights, FlashInfer, 32,768-token prefill chunks, 0.85 static-memory fraction, and the Qwen reasoning and tool parsers. The page publishes no speed result for that cell | Not a head-to-head with our 2x24GB NVFP4 probe. H200 FP8 uses a different checkpoint precision, memory class, kernel path, and prefill policy. It cannot replace our measured 57.32 / 96.22 t/s decode result or production's measured ~135 t/s. Winner on our deployed hardware remains Qwen3.6 | Recipe only: Qwen3.6 still wins measured speed |
| The official MTP overlay is EAGLE with 3 steps, top-k 1, and 4 draft tokens; small Blackwell guidance uses 2,048-token prefill chunks and warns that the default Mamba memory ratio can clamp concurrency | Our Phase 1 script already used the same EAGLE 3/1/4 flags, FP8 KV, and 2,048-token chunks, so the official page exposes no missed single-stream decode lever. We did not pin the Mamba ratio, but that governs state-pool concurrency rather than the one-request decode probe used for the promotion gate. No rerun or Phase 2 is justified | Core recipe independently matched |
The matched-budget repeats are done and they settle the quality axis cleanly. The composite is a config split: production +0.4 on one card (92.1 vs 91.7), candidate +1.2 on two (93.8 vs 92.6), candidate +0.4 on the mean - inside the ~2.6 noise band. The candidate's old tight-cap 93.8 was partly a budget artifact and its 89.8 entirely so. What survives matching: maths is a real candidate win (+5.2 mean, 97/100 vs 93.3/93.3), knowledge narrows to production +1.0, instruction-following goes to production +2.5, and token efficiency flips - production is ~16% leaner per scored item once caps match. The MTP opportunity flagged here is no longer hypothetical: measured on our own engine at +58% decode (see Source 11).
Standard intake chain first, then the extension arms that turn a screening result into a defended verdict.
The campaign is closed and the call is made. Qwen3.8-27B Q4_K_M fails all three promotion clauses: the quick-quality composite is a config split whose +0.4 mean sits inside the ~2.6 repeat-noise band (not a beat), the speed-weighted agentic /100 reads 71.1 vs 83.5 in production's favour, and the Q4 throughput gap - 3.4 to 4.1x slower at roughly 5x the energy per token - is a material speed regression by any reading of the bar. SGLang Phase 1 materially narrows that axis: NVFP4 MTP reaches 96.22 t/s, about 3x the llama.cpp Q4 serve, but it remains about 29% below production near 135 t/s and does not flip the speed clause. No clause is met; the bar requires all three. Production Qwen3.6-35B stays and SGLang Phase 2 stays off.
What makes this a lab-keeper rather than a discard: it is the first candidate to win a matched agentic cell against production (42 vs 41, thinking off, inside noise but at the measured suite ceiling - Opus 4.6 lands ≈41 on the same tasks with production's exact tier profile), it holds the best expert-tier score on file (16/18), the best tool-calling score on file (90.6 vs 84.4, +6.2 proven), a proven maths win (+5.2), and a working MTP speculative-decoding path worth +58% decode on our own engine. Operational flags if it is ever served: thinking must be OFF (the dose-response ladder 35/39/39/42 shows its own reasoning channel is the deficit), TC-SF-01 fired a destructive delete without confirming, and production stays ~16% token-leaner per scored item.
The campaign also paid for itself in methodology: card-bench has silently served q8_0 KV cache since inception (every published number is a q8-KV number - now a documented, measured default: true f16 KV is +10% decode / +21% prefill at 131K but cannot fit 131K on one 24 GB card), the kvq8 arm doubled as an exact-repro check and passed within 0.2 t/s, and the Q8_0 weight-quant arm confirmed Q4_K_M's discount is speed-positive (+56% decode at 8K) and ~10.5 GB lighter, not quality-funded. Production was restored at campaign end and verified live: both services active, health endpoint returning 200.
The agentic axis now has a three-configuration ladder, and the story turned. Thinking on at the matrix budget the candidate loses clearly: 35/45 vs production's 41/45, with every stall traced to hard truncation at the shared 4096-token cap. Raising the cap to 16K recovers half the gap (39/45 vs production's 42/45 at the same cap) but converts the worst task's truncation deaths into wall-clock timeout deaths - the thinking-token appetite is a real operational cost either way. Then the surprise: with thinking off entirely, at the original tight cap, the candidate scored 42/45 - zero truncations, zero loops, tying the best production number on file and beating the thinking-on matrix comparator. The agentic deficit is not "cannot do agentic work"; it is the cost of this model's own reasoning verbosity, and it vanishes when the reasoning channel is closed. That honesty flag is now closed: a production thinking-off cell was measured in the identical configuration and scored 41/45 - production gains nothing from turning thinking off (its fast MoE decode was never starved by the reasoning channel), so the matched cell reads candidate 42, production 41: the candidate's first outright win on the agentic axis. Two qualifiers stay attached: a one-point edge sits inside the repeat-noise band, and production's best number anywhere on file remains 42 (thinking on, 16K cap). The completed ladder's summary line: thinking costs this candidate 3 to 7 points; it costs production nothing. The matched-budget quick-quality repeats then settled the last caveated axes: the composite is a config split (production +0.4 on one card, candidate +1.2 on two, candidate +0.4 on the mean - inside noise), maths is a proven candidate win at +5.2, knowledge narrows to production +1.0, instruction-following goes to production +2.5, and token efficiency flips to production (~16% leaner per scored item; the earlier candidate lead was an output-cap artifact). Where that leaves the ledger: Qwen3.8 holds tool calling (+6.2, proven), maths (+5.2, proven), refusal behavior, VRAM footprint, and the matched agentic cell (42 vs 41, inside noise); Qwen3.6 holds every speed, power, and token-efficiency axis (proven, 3.4-4.4x on throughput) plus narrow leads on knowledge and instruction-following. The promotion bar is still not met: the agentic /100, which weighs speed, reads 71.1 vs 83.5 in production's favour. A frontier calibration pass reframes what the /45 numbers mean: Opus 4.6, driven through a first-party CLI harness over the same tasks, sandbox, and oracle, passed 17/19 tasks and failed exactly the two expert tasks production drops - at the same 1/3 rate on 3-run repeats. Full 3-run cells now exist for every expert and hard task, and they firm the picture: Opus lands expert 14/18 and hard 12/12 - production's tier profile to the point, one under the candidate's matched-cell 42. The suite's effective ceiling is 41-42/45, a frontier model sits on it alongside both finalists, and the /45 axis is saturated; what still separates the two local models is speed, power, and token efficiency, all of which production holds.
Opus 4.6 beats both local models outright on exactly two axes, and both are purple-marked in the head-to-head table: the quick-quality composite (94.4 vs the candidate's 92.75 and production's 92.35 config means) and knowledge (86.0 vs 80.0 and 81.0 - the axis where frontier-scale pretraining and its contamination exposure show up most directly). Everywhere else at least one Qwen matches or beats it. The candidate's maths mean of 98.5 edges Opus's 96.7. Instruction-following is a 95.0 tie with production. The long-context needle is a three-way 100. The refusal rate ties the candidate's best cell at 0.00. On the agentic suite Opus lands ≈41/45 - production's exact tier profile, one point under the candidate's matched-cell 42, on a suite whose measured ceiling is 41-42, so the /45 axis separates nobody at the top. Answered accuracy (0.36) is the one Opus number below the locals, and it stays flagged rather than concluded because the harness had no temperature control and a different system prompt. The overall read: at this suite's difficulty, a frontier model separates from well-served local 27-35B weights on knowledge breadth and a composite it drags up, and on nothing else that survives the noise band - while the locals run on two workstation cards at zero marginal cost per token.