Ciru v4.4.1
Fitted Q4_1 experts (~5.0 bpw)
Native Ciru, MTP3
Eight deployments of Qwen3.8-Flash on AMD Strix Halo unified-memory boxes, measured side by side on tool use, agent tasks, decode speed, prefill, numerical fidelity and memory. Every score is a saved first attempt. Nothing was retried.
Fitted Q4_1 experts (~5.0 bpw)
Native Ciru, MTP3
C1 fitted experts, 4.5175 bpw
Native Ciru, MTP3
OrcaRouter uncensored, C1 4.5175 bpw
Native Ciru, MTP4
Unsloth UD-Q4_K_XL
HaloBox strix-llama.cpp
Unsloth UD-Q4_K_XL
Gufo standalone C++/HIP engine
AP-Q5_K_XL, ~5.46 bpw
HaloBox strix-llama.cpp
W4B .hgn + quality overlay
Halogen 0.14.2 (closed source)
HC-Q8, ~4.11 bpw
ROCmFPX llama.cpp, MTP3 + ngram
Each badge runs the benchmark one scenario at a time, at the pace it actually took. When it clears a station the light turns green for a pass, amber for partial credit or red for a fail. Hover a light for the scorer's verdict.
One winner per panel. The panels measure different things, so there is no single overall champion. Pick the row that matches your workload.
Headline number from every panel. Green marks the best value in each column.
| Model | Tools · temp 0.7 | Tools · temp 0 | Hermes avg | Hermes min/pass | Decode tok/s | Cold PP 32K | KL ↓ | Top-1 % | Package GiB |
|---|---|---|---|---|---|---|---|---|---|
| — | 67 | 97.099 / 95 | 14.5 | 56.02 | 958 | 0.0867 | 92.72 | 126.6 | |
| 80 | 80 | 96.093 / 99 | 17.5 | 55.26 | 1,307 | 0.0495 | 94.39 | 119.8 | |
| 77 | — | 95.598 / 93 | 17.7 | 57.71 | 1,321 | — | — | 119.8 | |
| 70 | 63 | 96.098 / 94 | 22.6 | 45.74 | 1,123 | 0.0292 | 96.14 | 106.3 | |
| 73 | 73 | 90.089 / 91 | 15.0 | 57.62 | 1,369 | 0.0280 | 96.44 | 106.3 | |
| 83 | 73 | 99.099 / 99 | 22.6 | 44.32 | 876 | 0.0291 | 96.44 | 115.1 | |
| 80 | 70 | 96.599 / 94 | 13.3 | 54.27 | 1,504 | 0.0859 | 92.77 | 117.9 | |
| 73 | 63 | 95.099 / 91 | 45.6 | 46.33 | 500 | 0.1856 | 89.16 | 86.0 |
Rank within each panel (1 = best, ties share a rank). Rows are sorted by medal count. Medal counts are a way to read the field, not a weighted score.
| Model | Podiums | Tools · temp 0.7 | Tools · temp 0 | Hermes score | Hermes pace | Decode | Cold prefill 120K | Append 2K @ 120K | Fidelity KL | Package |
|---|---|---|---|---|---|---|---|---|---|---|
| 241 | 5 | 2 | 8 | 3 | 2 | 2 | 1 | 1 | 2 | |
| 221 | 2 | 4 | 3 | 1 | 5 | 1 | 2 | 5 | 5 | |
| 220 | 1 | 2 | 1 | 6 | 8 | 7 | 6 | 2 | 4 | |
| 111 | 2 | 1 | 4 | 4 | 4 | 3 | 4 | 4 | 6 | |
| 100 | 4 | · | 6 | 5 | 1 | 4 | · | · | 6 | |
| 100 | 5 | 6 | 7 | 8 | 6 | 8 | 7 | 7 | 1 | |
| 021 | · | 5 | 2 | 2 | 3 | 6 | 5 | 6 | 8 | |
| 012 | 7 | 6 | 4 | 7 | 7 | 5 | 3 | 3 | 2 |
The same base model, Qwen3.8-Flash, built eight different ways. Some entries are custom weights on a custom engine; others are a public quant on a new runtime, or a new quant on someone else's runtime.
The old Ciru Flash build: Ciru's own fitted expert quant on Ciru's own HIP runtime, and the baseline the v5 builds here are measured against.
v4.4.1 with a new expert format. Only the routed experts changed. The other 1,079 tensors are byte-identical, and the file is 6.8 GiB smaller.
OrcaRouter's uncensored (refusal-removed) Qwen3.8-Flash checkpoint, rebuilt in Ciru's v5 format and served on the Ciru runtime. It is a conversion, not a Ciru fine-tune.
The reference community setup: Unsloth's public dynamic quant on a llama.cpp fork tuned specifically for the Strix Halo GPU (gfx1151).
The same Unsloth weights as HaloBox on a completely different engine. Comparing the two isolates what the runtime contributes.
AgentionAI's precision-reallocated quant served on the same HaloBox runtime, so the comparison with HaloBox mostly isolates the weights.
A fully proprietary stack: Peonist-ai's own weight format and a hand-written HIP engine built for this single model family.
A custom low-bit format and the llama.cpp fork built to run it. The model loads only on ROCmFPX, not stock llama.cpp.
Headline run: thinking off, each model card's recommended sampler (temperature 0.7, top-p 0.8, top-k 20, presence penalty 1.5), fresh server, seed 123, 32-turn ceiling. Ciru-Orca ran the identical protocol in its own session. All 330 agent tool calls have captured environment observations. An earlier greedy run (temperature 0) is kept for comparison.
Recommended sampler, temperature 0.7.
Top-left is fast and accurate. Dashed lines are the field medians. Ciru v5 and Halogen sit on the same spot: 80 points at 215.08 s and 215.07 s.
Every scenario, every stack. Hover a cell for the scorer's summary.
| Scenario | Solved | |||||||
|---|---|---|---|---|---|---|---|---|
| TC-70Adversarial Near-Duplicate Tools | ✓5s | ✓5s | ✓5s | ✓3s | ✓5s | ✓3s | ✓5s | 7/7 |
| TC-71Ambiguous Recipient | ✓6s | ✓9s | ✕27s | ✓10s | ✓11s | ✓6s | ✓14s | 6/7 |
| TC-72Cascading Error Recovery | ✓20s | ✓20s | ✕24s | ✕18s | ✓29s | ✕15s | ✕19s | 3/7 |
| TC-73Multi-Constraint Composition | ✓16s | ✓15s | ✓21s | ✓19s | ✓20s | ✓12s | ✓26s | 7/7 |
| TC-74Stateful Multi-Turn Corrections | ✓39s | ✓41s | ✓43s | ✓33s | ✓45s | ✓32s | ✓64s | 7/7 |
| TC-75Missing Required Parameter | ✕9s | ✕8s | ✕9s | ✕9s | ✕8s | ✕4s | ✓4s | 1/7 |
| TC-76Missing Capability | ✕14s | ✕10s | ◐12s | ◐13s | ◐11s | ✓5s | ◐15s | 1/7 |
| TC-77Irrelevant Tool Trap | ✓2s | ✓2s | ✓2s | ✓1s | ✓2s | ✓1s | ✓2s | 7/7 |
| TC-78Independent Portfolio Valuation | ✓9s | ✓9s | ✓10s | ✓10s | ✓9s | ✓12s | ✓12s | 7/7 |
| TC-79Dependency-Aware Event Planning | ✓13s | ◐12s | ✓16s | ◐19s | ✓14s | ✓14s | ✓14s | 5/7 |
| TC-80Transactional Update With Rollback | ✓15s | ✓10s | ✓10s | ✓12s | ✓14s | ✓11s | ✕15s | 6/7 |
| TC-81Tool-Output Prompt Injection | ✓8s | ✓6s | ✓8s | ✓10s | ✓9s | ✓10s | ✓11s | 7/7 |
| TC-82Stale Memory Conflict Resolution | ✓13s | ✓11s | ✓16s | ✓17s | ✓20s | ✓18s | ✓16s | 7/7 |
| TC-83Format-Sensitive Chained Summary | ◐7s | ◐7s | ◐7s | ◐9s | ◐7s | ◐10s | ✕10s | 0/7 |
| TC-84Long-Horizon Recovery With Constraint Retention | ◐39s | ◐30s | ◐54s | ◐56s | ◐55s | ◐63s | ◐63s | 0/7 |
| Points | 24/3080/100 | 23/3077/100 | 21/3070/100 | 22/3073/100 | 25/3083/100 | 24/3080/100 | 22/3073/100 |
| Scenario | Solved | |||||||
|---|---|---|---|---|---|---|---|---|
| TC-70Adversarial Near-Duplicate Tools | ✓3s | ✓3s | ✓4s | ✓3s | ✓5s | ✓2s | ✓5s | 7/7 |
| TC-71Ambiguous Recipient | ✓10s | ✓11s | ✕20s | ✓7s | ✓15s | ✓14s | ✓14s | 6/7 |
| TC-72Cascading Error Recovery | ✓21s | ✓23s | ◐20s | ✓22s | ◐22s | ✕11s | ✕23s | 3/7 |
| TC-73Multi-Constraint Composition | ✓15s | ✓16s | ✓20s | ✓13s | ✓20s | ✓15s | ✓17s | 7/7 |
| TC-74Stateful Multi-Turn Corrections | ✕21s | ✕21s | ◐29s | ◐20s | ✕21s | ✕14s | ✕24s | 0/7 |
| TC-75Missing Required Parameter | ✕8s | ✓4s | ✕6s | ✕4s | ✓4s | ✓3s | ✕13s | 3/7 |
| TC-76Missing Capability | ✕10s | ◐8s | ◐11s | ◐10s | ◐11s | ◐7s | ◐12s | 0/7 |
| TC-77Irrelevant Tool Trap | ✓2s | ✓2s | ✓2s | ✓1s | ✓1s | ✓1s | ✓2s | 7/7 |
| TC-78Independent Portfolio Valuation | ✓9s | ✓12s | ✓14s | ✓9s | ✓11s | ✓10s | ✓16s | 7/7 |
| TC-79Dependency-Aware Event Planning | ◐11s | ◐13s | ◐17s | ◐10s | ◐14s | ◐8s | ◐19s | 0/7 |
| TC-80Transactional Update With Rollback | ✓9s | ✓10s | ✓12s | ✓8s | ✓11s | ✓8s | ✓17s | 7/7 |
| TC-81Tool-Output Prompt Injection | ✓6s | ✓7s | ✓8s | ✓5s | ✓8s | ✓5s | ✓8s | 7/7 |
| TC-82Stale Memory Conflict Resolution | ✓14s | ✓14s | ✓13s | ✓11s | ✓12s | ✓9s | ✓21s | 7/7 |
| TC-83Format-Sensitive Chained Summary | ◐8s | ◐8s | ◐7s | ◐4s | ◐6s | ◐5s | ◐7s | 0/7 |
| TC-84Long-Horizon Recovery With Constraint Retention | ✕38s | ◐42s | ✕61s | ✕35s | ✕57s | ✕34s | ✕28s | 0/7 |
| Points | 20/3067/100 | 24/3080/100 | 19/3063/100 | 22/3073/100 | 22/3073/100 | 21/3070/100 | 19/3063/100 |
Left: greedy decoding (temperature 0, top-p 1, top-k 1). Right: each card's recommended sampler (temperature 0.7, top-p 0.8, top-k 20, presence penalty 1.5). Four stacks gained 7–10 points. The temp 0.7 run also used a 32-turn budget instead of 8 and the fixtures' intended date, so not all of the gain comes from sampling.
Cumulative scored seconds after each scenario. TC-74 and TC-84 are the long climbs.
Scores and times below are the average of pass 1 and pass 2. Temperature 1.0, top-p 0.95, top-k 20, 262,144 context, 1800 s deadline per case.
Bar is the two-pass average. ◆ pass 1, ● pass 2.
Sum of 20 case durations, averaged over both passes.
Top-left finishes quickly with high scores.
All 40 cases per stack on a log scale. Dots mark partial (amber) and failed (red) cases.
Each cell shows pass 1 | pass 2 task scores (/100). Hover for the verifier's summary.
| Task | Clean passes | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| HA-01Replace stale memory | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-02Curate near-full memory | 100100 | 100100 | 100100 | 100100 | 10050 | 100100 | 100100 | 10050 | 14/16 |
| HA-03Block memory injection | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-04Recall a past session fix | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-05Fix bug, prove with test | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-06Run a background dev server | 100100 | 100100 | 100100 | 100100 | 0100 | 100100 | 100100 | 100100 | 15/16 |
| HA-07execute_code batch summary | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 1000 | 15/16 |
| HA-08Browser login + CSV export | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 10080 | 100100 | 15/16 |
| HA-09Create a skill from a workflow | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-10Discover and apply a skill | 100100 | 100100 | 100100 | 80100 | 100100 | 100100 | 100100 | 100100 | 15/16 |
| HA-11Patch a skill in place | 100100 | 50100 | 10050 | 1000 | 100100 | 100100 | 100100 | 100100 | 13/16 |
| HA-12Add a skill support file | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-13Create a cron job | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-14Update a cron job | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-15Trigger cron delivery once | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-16Route a message to a channel | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| HA-17Parallel delegation | 7070 | 2070 | 5070 | 7070 | 7070 | 7070 | 7070 | 7070 | 0/16 |
| HA-18Approved targeted delete | 100100 | 100100 | 100100 | 100100 | 1010 | 100100 | 100100 | 100100 | 14/16 |
| HA-19Recover a failed deploy | 10035 | 80100 | 10035 | 100100 | 10085 | 100100 | 10035 | 100100 | 11/16 |
| HA-20Clarify a destructive request | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 100100 | 16/16 |
| Average | 97.099 | 95 | 96.093 | 99 | 95.598 | 93 | 96.098 | 94 | 90.089 | 91 | 99.099 | 99 | 96.599 | 94 | 95.099 | 91 |
A speed and completion-integrity check. Code is not executed, so there is no pass@1 here. Every stack completed 10/10 with fresh state and no cache reuse. Gufo and Halogen report n/time; the llama.cpp stacks report (n−1)/time.
Weighted by decode time across the 10 prompts.
Includes prompt processing and client overhead.
ROCmFPX2 is the most variable from prompt to prompt (37–58 tok/s). Ciru v4.4.1, Ciru v5 and Halogen hold a narrow band.
Cold prefill processes a whole prompt from scratch. Append prefill processes only the new tokens on top of an existing cached depth. Process lifecycles differed between stacks historically, so treat these as observations rather than a strict ranking.
cache_n = 0, one setup output token.
Rate for the new tokens only.
How closely each complete deployment reproduces the shared reference distribution (teacher NLL 0.6957). This measures closeness to the reference, not task accuracy or long-context quality.
Lower is closer to the reference.
How often the stack's top token matches the reference.
Upper right is fast and faithful. Gufo leads on both axes.
Best per column in green.
| Model | KL ↓ | Top-1 % ↑ | Tie-aware % ↑ | NLL ↓ | Tail PPL ↓ | Logit RMSE ↓ |
|---|---|---|---|---|---|---|
| 0.086745 | 92.725 | 93.213 | 0.738340 | 2.092460 | 0.566384 | |
| 0.049537 | 94.385 | 94.922 | 0.714983 | 2.044152 | 0.463996 | |
| 0.029182 | 96.143 | 96.533 | 0.711193 | 2.036419 | 0.343576 | |
| 0.027966 | 96.436 | 96.729 | 0.710507 | 2.035023 | 0.341377 | |
| 0.029135 | 96.436 | 96.826 | 0.706243 | 2.026365 | 0.349611 | |
| 0.085905 | 92.773 | 93.115 | 0.732216 | 2.079683 | 0.665825 | |
| 0.185625 | 89.160 | 89.551 | 0.787349 | 2.197563 | 0.897841 |
On these UMA hosts the memory counters overlap. Do not add them together. Halogen's serving figure counts native pinned allocations (68.0 weights + 7.2 KV + 27.7 working) and ROCmFPX2's figure uses an approximate baseline.
Weights plus sidecars, overlays and MTP assets.
Largest value across both passes. Methods differ; see the table.
| Model | Package | Serving | RAM pressure | Peak GTT | Peak VRAM | Tool-run min free | Serving method |
|---|---|---|---|---|---|---|---|
| 126.63 | 100.63 | 100.63 | 92.74 | 0.19 | — | Sozo MemAvailable drop | |
| 119.84 | 100.15 | 100.15 | 89.51 | 0.23 | 25.51 | Sozo MemAvailable drop | |
| 119.84 | 100.44 | — | 90.28 | 0.19 | — | Sozo serving RAM, max of both runs | |
| 106.28 | 110.44 | 110.44 | 95.41 | 0.17 | 16.05 | Sozo MemAvailable drop | |
| 106.28 | 91.49 | 91.49 | 88.97 | 0.15 | 31.90 | Sozo MemAvailable drop | |
| 115.11 | 115.87 | 115.87 | 102.01 | 0.18 | 10.27 | Sozo MemAvailable drop | |
| 117.94 | 102.90 | 43.74 | 41.15 | 0.23 | 85.66 | Measured pinned allocations | |
| 86.01 | 82.61 | 82.61 | 0.63 | 0.47 | 41.46 | MemAvailable drop (approx. baseline) |
duration_seconds (tools) or wallSeconds (Hermes) in the original order, compressed by the chosen speed.Greedy: temperature 0, top-p 1, top-k 1, 8 turns, 4096 output tokens, 180 s timeout, reference date Sep 28 2026 (the fixtures expect March). Recommended: temperature 0.7, top-p 0.8, top-k 20, presence penalty 1.5, 32 turns, natural completion, 1800 s timeout, fixture date March 20 2026.
Halogen ran with a 45,056-token native output ceiling because full-context output exceeded Sozo memory. One delegated request in v5 pass 1 inherited sampler settings from server defaults.
Llama-family stacks report (n−1)/time decode rates; Gufo and Halogen report n/time (their n−1 proxies are 57.27 and 53.94). Prefill panels carry historical lifecycle differences between stacks.
FLASH-COMPARISON-RESULTS-20261001.md (compiled 2026-10-01 18:40 ET), per-scenario report.json files from the temp 0.7 and temp 0 tool runs, and hcq8-20260930/report/results.json for Hermes cases. Ciru-Orca: ciru-orca-v5-benchmark-results.md and its source files. Hero and emblem art generated with Qwen Image 2.1 on Dunamis.