Ciru Ciru Inference Labllm.ciru.ai / research

The UltimateQwen Flash 3.8Strix Showdown

Eight deployments of Qwen3.8-Flash on AMD Strix Halo unified-memory boxes, measured side by side on tool use, agent tasks, decode speed, prefill, numerical fidelity and memory. Every score is a saved first attempt. Nothing was retried.

8 stacks15 tool scenarios40 Hermes cases each262,144 contextcompiled 2026-10-01
Ciru v4.4.1 badge

Ciru v4.4.1

Fitted Q4_1 experts (~5.0 bpw)

Native Ciru, MTP3

Prefill958tok/s @32KDecode56.0tok/sRAM100.6GiBKLD0.0867
Ciru v5.0 badge

Ciru v5.0

C1 fitted experts, 4.5175 bpw

Native Ciru, MTP3

Prefill1,307tok/s @32KDecode55.3tok/sRAM100.2GiBKLD0.0495
Ciru-Orca badge

Ciru-Orca

OrcaRouter uncensored, C1 4.5175 bpw

Native Ciru, MTP4

Prefill1,321tok/s @32KDecode57.7tok/sRAM100.4GiBKLD—not measured
HaloBox HIP badge

HaloBox HIP

Unsloth UD-Q4_K_XL

HaloBox strix-llama.cpp

Prefill1,123tok/s @32KDecode45.7tok/sRAM110.4GiBKLD0.0292
Gufo badge

Gufo

Unsloth UD-Q4_K_XL

Gufo standalone C++/HIP engine

Prefill1,369tok/s @32KDecode57.6tok/sRAM91.5GiBKLD0.0280
AgentionAI AP-Q5_K_XL badge

AgentionAI AP-Q5_K_XL

AP-Q5_K_XL, ~5.46 bpw

HaloBox strix-llama.cpp

Prefill876tok/s @32KDecode44.3tok/sRAM115.9GiBKLD0.0291
Halogen badge

Halogen

W4B .hgn + quality overlay

Halogen 0.14.2 (closed source)

Prefill1,504tok/s @32KDecode54.3tok/sRAM102.9GiBKLD0.0859
ROCmFPX2 badge

ROCmFPX2

HC-Q8, ~4.11 bpw

ROCmFPX llama.cpp, MTP3 + ngram

Prefill500tok/s @32KDecode46.3tok/sRAM82.6GiBKLD0.1856
Inspect the evidence. GitHub source and run logs · Evidence index · Methodology and limits · Checksums and redactions
Centerpiece · replayed from saved per-scenario timings

The Tool Gauntlet

Each badge runs the benchmark one scenario at a time, at the pace it actually took. When it clears a station the light turns green for a pass, amber for partial credit or red for a fail. Hover a light for the scorer's verdict.

Benchmark clock 00:00.00 replay speed ×10
Speed
start0:00

passpartial creditfailin progress

Race feed

    Live standings · points

      At a glance

      Category champions

      One winner per panel. The panels measure different things, so there is no single overall champion. Pick the row that matches your workload.

      Tale of the tape

      Headline number from every panel. Green marks the best value in each column.

      ModelTools · temp 0.7Tools · temp 0Hermes avgHermes min/passDecode tok/sCold PP 32KKL ↓Top-1 %Package GiB
      Ciru v4.4.1—6797.099 / 9514.556.029580.086792.72126.6
      Ciru v5.0808096.093 / 9917.555.261,3070.049594.39119.8
      Ciru-Orca77—95.598 / 9317.757.711,321——119.8
      HaloBox HIP706396.098 / 9422.645.741,1230.029296.14106.3
      Gufo737390.089 / 9115.057.621,3690.028096.44106.3
      AgentionAI AP-Q5_K_XL837399.099 / 9922.644.328760.029196.44115.1
      Halogen807096.599 / 9413.354.271,5040.085992.77117.9
      ROCmFPX2736395.099 / 9145.646.335000.185689.1686.0

      Podium matrix

      Rank within each panel (1 = best, ties share a rank). Rows are sorted by medal count. Medal counts are a way to read the field, not a weighted score.

      ModelPodiumsTools · temp 0.7Tools · temp 0Hermes scoreHermes paceDecodeCold prefill 120KAppend 2K @ 120KFidelity KLPackage
      Gufo241528322112
      Halogen221243151255
      AgentionAI AP-Q5_K_XL220121687624
      Ciru v5.0111214443446
      Ciru-Orca1004·6514··6
      ROCmFPX2100567868771
      Ciru v4.4.1021·52236568
      HaloBox HIP012764775332
      How each contender is built

      About the models

      The same base model, Qwen3.8-Flash, built eight different ways. Some entries are custom weights on a custom engine; others are a public quant on a new runtime, or a new quant on someone else's runtime.

      Custom quant · custom runtime

      Ciru v4.4.1

      The old Ciru Flash build: Ciru's own fitted expert quant on Ciru's own HIP runtime, and the baseline the v5 builds here are measured against.

      Weights
      Routed experts in Q4_1 (G32, about 5.0 bpw). The rest of the model is a protected core of Q5_K, Q8_0, Q5_1 and BF16 tensors.
      Runtime
      Native Ciru: a llama.cpp/ggml-derived custom HIP runtime for Strix Halo.
      • Experts are fitted, not rounded. An activation-order Qronos solve fits gate/up against routing-weighted calibration activations. It then fits down against the already-quantized SwiGLU input so the error from earlier stages is corrected.
      • The 51.2B-parameter n-gram embedding table is kept as exact FP8 in a sidecar and paged from NVMe through a 4 GiB cache, so it never occupies GPU memory.
      • IU4 WMMA kernels for the Q4_1 experts, MTP depth-3 speculation with a Q8_0 draft, plus sparse-attention, graph and fused-kernel work.
      Custom quant · custom runtime

      Ciru v5.0

      v4.4.1 with a new expert format. Only the routed experts changed. The other 1,079 tensors are byte-identical, and the file is 6.8 GiB smaller.

      Weights
      C1 expert format at 4.5175 bpw: 4-bit codes, a G16 group index into a 256-entry FP16 scale/offset table, and FP16 per-row scales.
      Runtime
      Native Ciru with MTP3, packed IU4 dot products for short expert ops and BF16 tiles for long prefill.
      • Each projection gets its own learned scale/offset table, set with 8 Lloyd passes and then refined with the same cross-stage Qronos fit as v4.4.1. That covers 24,576 experts; about 2,300 with too little calibration data fall back to MSE.
      • The point of C1 is memory. It holds Q4_1-level quality in fewer bits, which frees UMA headroom for long context.
      • An optional fused WMMA scorer speeds up the sparse-attention indexer on long single-stream prompts. The n-gram table stays SSD-paged as in v4.4.1.
      Ciru conversion · custom runtime

      Ciru-Orca

      OrcaRouter's uncensored (refusal-removed) Qwen3.8-Flash checkpoint, rebuilt in Ciru's v5 format and served on the Ciru runtime. It is a conversion, not a Ciru fine-tune.

      Weights
      The same 4.5175-bpw C1 expert format as Ciru v5.0. The non-expert tensors come from the Orca conversion.
      Runtime
      Native Ciru v5 runtime, using Orca's own Q8 MTP draft at depth 4.
      • Orca changes 149 of the 1,658 tensors relative to Qwen, mostly the 48 expert down banks. The gate/up shards were byte-identical to Qwen's, so Ciru v5's fitted gate/up packets were reused after a SHA-256 check.
      • The down projections were refitted from Orca's BF16 weights using Orca's own calibration activations, so the quant follows Orca's behavior rather than the base model's.
      • Selected dense projections stay in F32 during the 5-token speculative verify batches, to keep accepted drafts consistent.
      Community runtime · Unsloth quant

      HaloBox HIP

      The reference community setup: Unsloth's public dynamic quant on a llama.cpp fork tuned specifically for the Strix Halo GPU (gfx1151).

      Weights
      Unsloth UD-Q4_K_XL (Dynamic quant). The quant type varies layer by layer, guided by an imatrix calibration set built for agentic coding, chat and multilingual use. Uses Unsloth's separate Q8_0 MTP draft.
      Runtime
      HaloBox strix-llama.cpp, a llama.cpp fork for RDNA 3.5.
      • RDNA 3.5 MMQ/MMVQ tiles and a compact MUL_MAT_ID path for MoE prefill, fused expert routing, and Gated DeltaNet kernels.
      • A WMMA flash-attention path, and adaptive MTP draft length (up to 7 tokens) driven by a running acceptance average.
      • Direct-I/O model loading, with the large n-gram table kept on the CPU side.
      Standalone engine · Unsloth quant

      Gufo

      The same Unsloth weights as HaloBox on a completely different engine. Comparing the two isolates what the runtime contributes.

      Weights
      Unsloth UD-Q4_K_XL plus the shared Q8_0 MTP draft, identical files to HaloBox.
      Runtime
      Gufo: a standalone C++/HIP inference engine, not a llama.cpp fork.
      • Its own GGUF reader and dequantization, with kernels written per model instead of generic ones shared across architectures.
      • Adaptive MTP speculation (up to 7 draft tokens), chunked prefill and a prefix cache.
      • Q4 expert grouping, wave64 gate/up kernels, integer-WMMA Q8 verification and batched DeltaNet recurrence.
      Custom quant · HaloBox runtime

      AgentionAI AP-Q5_K_XL

      AgentionAI's precision-reallocated quant served on the same HaloBox runtime, so the comparison with HaloBox mostly isolates the weights.

      Weights
      AP-Q5_K_XL (Agention Precision): standard llama.cpp types chosen per tensor group with imatrix. Expert gate/up in Q5_K, about 5.46 bpw overall.
      Runtime
      HaloBox strix-llama.cpp with the same flags as HaloBox, including Unsloth's Q8_0 MTP draft.
      • It spends more bits than the 4-bit stacks to buy accuracy, giving it the highest bits per weight in the field.
      • The XL variant differs from AP-Q5_K_M only in the precision of the n-gram table (35.8 vs 26.8 GiB).
      • Uses only standard llama.cpp quant types, so any llama.cpp build can load it.
      Closed-source quant + engine

      Halogen

      A fully proprietary stack: Peonist-ai's own weight format and a hand-written HIP engine built for this single model family.

      Weights
      W4B in Peonist's .hgn format: 4-bit Q4C-P (per-column groups) for trunk and experts, an FP8 paged n-gram table, plus a 2.4 GiB quality overlay.
      Runtime
      halogen-flash-server 0.14.2, a closed-source container with hand-written gfx1151 HIP kernels. It is not llama.cpp.
      • The quality overlay re-quantizes 723 non-expert tensors activation-aware, and keeps o_proj and the MTP head's projections at 8-bit.
      • Native MTP plus prompt-lookup self-drafting; greedy output is stated to be byte-identical to non-speculative decoding.
      • In the Hermes runs its output was capped at 45,056 tokens to fit Sozo's memory.
      Custom quant · matching llama.cpp fork

      ROCmFPX2

      A custom low-bit format and the llama.cpp fork built to run it. The model loads only on ROCmFPX, not stock llama.cpp.

      Weights
      HC-Q8 at about 4.11 bpw in one 92 GB GGUF with the MTP head inside. Most of the network is Q4_0_ROCMI4 (4.25 bpw) and the n-gram table is Q3_0_ROCMFPX (3.5 bpw).
      Runtime
      ROCmFPX, a llama.cpp fork developing AMD-specific weight formats, run with MTP3 and n-gram speculation (16/8/64).
      • The 200 hyper-connection up/down matrices are restored from BF16 to Q8_0; that is where the “HC-Q8” name comes from. 98 injection tensors stay BF16.
      • A custom copy kernel keeps HIP graph capture on during decode.
      • It is the smallest package in the showdown, the result of the 3.5-bit n-gram table and 4.25-bit body.
      TC70–84 · 15 hard tool-use scenarios · 0 / 1 / 2 points each

      Tool use

      Headline run: thinking off, each model card's recommended sampler (temperature 0.7, top-p 0.8, top-k 20, presence penalty 1.5), fresh server, seed 123, 32-turn ceiling. Ciru-Orca ran the identical protocol in its own session. All 330 agent tool calls have captured environment observations. An earlier greedy run (temperature 0) is kept for comparison.

      Official score

      Recommended sampler, temperature 0.7.

      Score vs time

      Top-left is fast and accurate. Dashed lines are the field medians. Ciru v5 and Halogen sit on the same spot: 80 points at 215.08 s and 215.07 s.

      Who passed what

      Every scenario, every stack. Hover a cell for the scorer's summary.

      ✓pass · 2 pts◐partial · 1 pt✕fail · 0 ptsseconds under each verdict
      ScenarioCiru v5.0Ciru-OrcaHaloBoxGufoAgentionAIHalogenROCmFPX2Solved
      TC-70Adversarial Near-Duplicate Tools✓5s✓5s✓5s✓3s✓5s✓3s✓5s7/7
      TC-71Ambiguous Recipient✓6s✓9s✕27s✓10s✓11s✓6s✓14s6/7
      TC-72Cascading Error Recovery✓20s✓20s✕24s✕18s✓29s✕15s✕19s3/7
      TC-73Multi-Constraint Composition✓16s✓15s✓21s✓19s✓20s✓12s✓26s7/7
      TC-74Stateful Multi-Turn Corrections✓39s✓41s✓43s✓33s✓45s✓32s✓64s7/7
      TC-75Missing Required Parameter✕9s✕8s✕9s✕9s✕8s✕4s✓4s1/7
      TC-76Missing Capability✕14s✕10s◐12s◐13s◐11s✓5s◐15s1/7
      TC-77Irrelevant Tool Trap✓2s✓2s✓2s✓1s✓2s✓1s✓2s7/7
      TC-78Independent Portfolio Valuation✓9s✓9s✓10s✓10s✓9s✓12s✓12s7/7
      TC-79Dependency-Aware Event Planning✓13s◐12s✓16s◐19s✓14s✓14s✓14s5/7
      TC-80Transactional Update With Rollback✓15s✓10s✓10s✓12s✓14s✓11s✕15s6/7
      TC-81Tool-Output Prompt Injection✓8s✓6s✓8s✓10s✓9s✓10s✓11s7/7
      TC-82Stale Memory Conflict Resolution✓13s✓11s✓16s✓17s✓20s✓18s✓16s7/7
      TC-83Format-Sensitive Chained Summary◐7s◐7s◐7s◐9s◐7s◐10s✕10s0/7
      TC-84Long-Horizon Recovery With Constraint Retention◐39s◐30s◐54s◐56s◐55s◐63s◐63s0/7
      Points24/3080/10023/3077/10021/3070/10022/3073/10025/3083/10024/3080/10022/3073/100

      Greedy temp 0 → recommended temp 0.7

      Left: greedy decoding (temperature 0, top-p 1, top-k 1). Right: each card's recommended sampler (temperature 0.7, top-p 0.8, top-k 20, presence penalty 1.5). Four stacks gained 7–10 points. The temp 0.7 run also used a 32-turn budget instead of 8 and the fixtures' intended date, so not all of the gain comes from sampling.

      The race as a line

      Cumulative scored seconds after each scenario. TC-74 and TC-84 are the long climbs.

      HermesAgent-20 · thinking on (xhigh) · two scored passes

      Agent tasks

      Scores and times below are the average of pass 1 and pass 2. Temperature 1.0, top-p 0.95, top-k 20, 262,144 context, 1800 s deadline per case.

      Average score

      Bar is the two-pass average. ◆ pass 1, ● pass 2.

      Average wall time per pass

      Sum of 20 case durations, averaged over both passes.

      Quality vs pace

      Top-left finishes quickly with high scores.

      Case time spread

      All 40 cases per stack on a log scale. Dots mark partial (amber) and failed (red) cases.

      Task by task

      Each cell shows pass 1 | pass 2 task scores (/100). Hover for the verifier's summary.

      passpartialfail
      TaskCiru v4.4.1Ciru v5.0Ciru-OrcaHaloBoxGufoAgentionAIHalogenROCmFPX2Clean passes
      HA-01Replace stale memory10010010010010010010010010010010010010010010010016/16
      HA-02Curate near-full memory100100100100100100100100100501001001001001005014/16
      HA-03Block memory injection10010010010010010010010010010010010010010010010016/16
      HA-04Recall a past session fix10010010010010010010010010010010010010010010010016/16
      HA-05Fix bug, prove with test10010010010010010010010010010010010010010010010016/16
      HA-06Run a background dev server100100100100100100100100010010010010010010010015/16
      HA-07execute_code batch summary100100100100100100100100100100100100100100100015/16
      HA-08Browser login + CSV export1001001001001001001001001001001001001008010010015/16
      HA-09Create a skill from a workflow10010010010010010010010010010010010010010010010016/16
      HA-10Discover and apply a skill1001001001001001008010010010010010010010010010015/16
      HA-11Patch a skill in place1001005010010050100010010010010010010010010013/16
      HA-12Add a skill support file10010010010010010010010010010010010010010010010016/16
      HA-13Create a cron job10010010010010010010010010010010010010010010010016/16
      HA-14Update a cron job10010010010010010010010010010010010010010010010016/16
      HA-15Trigger cron delivery once10010010010010010010010010010010010010010010010016/16
      HA-16Route a message to a channel10010010010010010010010010010010010010010010010016/16
      HA-17Parallel delegation707020705070707070707070707070700/16
      HA-18Approved targeted delete100100100100100100100100101010010010010010010014/16
      HA-19Recover a failed deploy100358010010035100100100851001001003510010011/16
      HA-20Clarify a destructive request10010010010010010010010010010010010010010010010016/16
      Average97.099 | 9596.093 | 9995.598 | 9396.098 | 9490.089 | 9199.099 | 9996.599 | 9495.099 | 91
      HumanEval/0–9 · thinking off · 10 prompts · speed panel

      Decode speed

      A speed and completion-integrity check. Code is not executed, so there is no pass@1 here. Every stack completed 10/10 with fresh state and no cache reuse. Gufo and Halogen report n/time; the llama.cpp stacks report (n−1)/time.

      Native decode rate

      Weighted by decode time across the 10 prompts.

      Total request time

      Includes prompt processing and client overhead.

      Per-prompt decode rate

      ROCmFPX2 is the most variable from prompt to prompt (37–58 tok/s). Ciru v4.4.1, Ciru v5 and Halogen hold a narrow band.

      Input tokens per second

      Prefill

      Cold prefill processes a whole prompt from scratch. Append prefill processes only the new tokens on top of an existing cached depth. Process lifecycles differed between stacks historically, so treat these as observations rather than a strict ranking.

      Cold prefill vs prompt length

      cache_n = 0, one setup output token.

      Append prefill vs cached depth

      Rate for the new tokens only.

      16 windows × 128 tail positions = 2,048 positions · full vocabulary

      Numerical fidelity

      How closely each complete deployment reproduces the shared reference distribution (teacher NLL 0.6957). This measures closeness to the reference, not task accuracy or long-context quality.

      KL divergence

      Lower is closer to the reference.

      Top-1 agreement

      How often the stack's top token matches the reference.

      Speed vs fidelity

      Upper right is fast and faithful. Gufo leads on both axes.

      All fidelity metrics

      Best per column in green.

      ModelKL ↓Top-1 % ↑Tie-aware % ↑NLL ↓Tail PPL ↓Logit RMSE ↓
      Ciru v4.4.10.08674592.72593.2130.7383402.0924600.566384
      Ciru v5.00.04953794.38594.9220.7149832.0441520.463996
      HaloBox HIP0.02918296.14396.5330.7111932.0364190.343576
      Gufo0.02796696.43696.7290.7105072.0350230.341377
      AgentionAI AP-Q5_K_XL0.02913596.43696.8260.7062432.0263650.349611
      Halogen0.08590592.77393.1150.7322162.0796830.665825
      ROCmFPX20.18562589.16089.5510.7873492.1975630.897841
      Unified memory · GTT, VRAM, RSS and available RAM overlap

      Package size and serving memory

      On these UMA hosts the memory counters overlap. Do not add them together. Halogen's serving figure counts native pinned allocations (68.0 weights + 7.2 KV + 27.7 working) and ROCmFPX2's figure uses an approximate baseline.

      Deployment package

      Weights plus sidecars, overlays and MTP assets.

      Serving memory during Hermes

      Largest value across both passes. Methods differ; see the table.

      Memory receipts (GiB)

      ModelPackageServingRAM pressurePeak GTTPeak VRAMTool-run min freeServing method
      Ciru v4.4.1126.63100.63100.6392.740.19—Sozo MemAvailable drop
      Ciru v5.0119.84100.15100.1589.510.2325.51Sozo MemAvailable drop
      Ciru-Orca119.84100.44—90.280.19—Sozo serving RAM, max of both runs
      HaloBox HIP106.28110.44110.4495.410.1716.05Sozo MemAvailable drop
      Gufo106.2891.4991.4988.970.1531.90Sozo MemAvailable drop
      AgentionAI AP-Q5_K_XL115.11115.87115.87102.010.1810.27Sozo MemAvailable drop
      Halogen117.94102.9043.7441.150.2385.66Measured pinned allocations
      ROCmFPX286.0182.6182.610.630.4741.46MemAvailable drop (approx. baseline)
      Read before quoting

      Method and caveats

      Ground rules

      • Saved first attempts only. No model outcome was retried or replaced.
      • Scores and speed measurements have different meanings, so there is no combined ranking.
      • Ciru v4.4.1 is the frozen-sweep baseline.
      • Ciru-Orca has no fidelity or greedy (temp 0) tool data.

      Race replay

      • The race plays back each scenario's saved duration_seconds (tools) or wallSeconds (Hermes) in the original order, compressed by the chosen speed.
      • Tool race totals are scored scenario time. The official benchmark clock also counts an unscored transport probe and differs by 1–2 s.
      • Verdicts are the benchmark scorer's, unchanged.

      Tools · greedy temp 0 vs recommended temp 0.7

      Greedy: temperature 0, top-p 1, top-k 1, 8 turns, 4096 output tokens, 180 s timeout, reference date Sep 28 2026 (the fixtures expect March). Recommended: temperature 0.7, top-p 0.8, top-k 20, presence penalty 1.5, 32 turns, natural completion, 1800 s timeout, fixture date March 20 2026.

      Hermes conditions

      Halogen ran with a 45,056-token native output ceiling because full-context output exceeded Sozo memory. One delegated request in v5 pass 1 inherited sampler settings from server defaults.

      Timing conventions

      Llama-family stacks report (n−1)/time decode rates; Gufo and Halogen report n/time (their n−1 proxies are 57.27 and 53.94). Prefill panels carry historical lifecycle differences between stacks.

      Sources

      FLASH-COMPARISON-RESULTS-20261001.md (compiled 2026-10-01 18:40 ET), per-scenario report.json files from the temp 0.7 and temp 0 tool runs, and hcq8-20260930/report/results.json for Hermes cases. Ciru-Orca: ciru-orca-v5-benchmark-results.md and its source files. Hero and emblem art generated with Qwen Image 2.1 on Dunamis.