Aggregate generation rate
Server decode time · higher is faster
CIRU v3 on Strix Halo: faster long-context serving, measured task completion times, and agent-quality checks. Compare the previous CIRU runner and Halo, then explore the retained v2 evidence.
Same CIRU weights, updated runtime and serving profile. At 64K, v3 reduces whole-request time by 23.93%, raises prompt throughput by 29.99%, and raises generation throughput by 81.66% versus the previous CIRU runner. At 4K, whole-request time falls 11.80%.
All three arms ran sequentially on Ciru: Ryzen AI Max+ 395 / gfx1151, 128 GB unified memory, NixOS. Identical input token IDs, cold prompt cache, 128 output tokens, one slot and 262,144 context capacity. Each load has an excluded warmup. Previous CIRU is the locally qualified v2.0.1 runner; it was not separately published as a tag.
| Input tokens | Profile | Prompt tok/s | Generation tok/s | First piece · s | Whole request · s | Requests |
|---|---|---|---|---|---|---|
| 4,096 | Previous CIRU | 392.00 | 22.52 | 10.70 | 16.34 | 2 |
| 4,096 | CIRU v3 | 455.65 | 24.60 | 9.25 | 14.41 | 6 |
| 4,096 | Halo | 381.49 | 35.30 | 11.09 | 14.69 | 4 |
| 65,536 | Previous CIRU | 284.49 | 13.33 | 230.46 | 239.99 | 1 |
| 65,536 | CIRU v3 | 369.81 | 24.22 | 177.32 | 182.57 | 3 |
| 65,536 | Halo | 263.42 | 23.28 | 248.91 | 254.37 | 2 |
Seconds · lower is faster
Minutes · includes tools and scoring
Previous CIRU uses MTP6, b2048/u512; v3 uses MTP6, b1024/u1024. Both use F16 target KV and Q8 draft KV. Halo uses its unmodified Vulkan runner, Unsloth UD-Q4_K_XL target, EasiiX Q8 head and screened MTP3 profile. Halo generates faster at 4K; the small 4K whole-request difference is not a robust general v3 win. Three v3 loads, two Halo loads and one previous-CIRU load do not establish confidence intervals. This measures complete serving profiles, not runtime changes in isolation.
Nonthinking, temperature 0.7, top-p 0.8, top-k 20, min-p 0, presence penalty 1.5, repeat penalty 1, frequency penalty 0, seed 123. EOS is honored; all reported requests produced 128 tokens. Native TG excludes the first output token; first-piece latency is the first streamed content-field event.
| Input | Profile | Peak system RAM · GiB | Max RAM increase · GiB | Peak GTT · GiB | Peak VRAM · GiB |
|---|---|---|---|---|---|
| 4,096 | Previous CIRU | 103.43 | 2.15 | 88.58 | 0.52 |
| 65,536 | Previous CIRU | 101.32 | 0.72 | 88.58 | 0.52 |
| 4,096 | CIRU v3 | 106.02 | 2.43 | 90.29 | 0.53 |
| 65,536 | CIRU v3 | 103.66 | 0.74 | 90.29 | 0.52 |
| 4,096 | Halo | 104.88 | 2.77 | 93.08 | 1.85 |
| 65,536 | Halo | 103.72 | 0.60 | 93.17 | 1.99 |
The v3 capacity check completed 261,888 input + 128 output tokens: 257.44 prompt tok/s, 18.00 generation tok/s, 1017.38 s to first piece and 1024.44 s whole request. One CIRU-only request establishes serving capacity, not full-context answer accuracy or a 256K comparison with Halo.
The additional Ornith difficulty panel combines 22 short IFEval/GSM8K/HumanEval tasks, six Hermes scenarios repeated twice, eight hard tasks with a shared 63K-token history, and short/long coding health checks. It uses cases selected from earlier disagreements and failures; scores do not estimate general benchmark accuracy.
| Stage | Previous CIRU | CIRU v3 | Halo MTP3 |
|---|---|---|---|
| Complete panel, after readiness | 29m 17s | 25m 18s | 24m 34s |
| Model load, additional | 0m 31s | 0m 31s | 0m 34s |
| Short scored stage | 5m 20s | 5m 21s | 4m 36s |
| Hermes, two rounds combined | 12m 28s | 11m 38s | 10m 08s |
| 63K history seeding | 3m 26s | 2m 43s | 3m 21s |
| Long hard stage, history already loaded | 4m 44s | 2m 42s | 3m 23s |
| All generated tokens | 33,294 | 33,451 | 31,058 |
V3 takes 13.59% less total time than previous CIRU, while Halo finishes 43.89 seconds sooner than v3. Long hard-stage time falls 42.98% versus previous CIRU, but output lengths differ. Short-stage time is effectively unchanged. These are complete workload timings, not equal-output generation speedups.
| Profile | Short IFEval | Short GSM8K | Short HumanEval | Long IFEval | Long GSM8K | Long HumanEval |
|---|---|---|---|---|---|---|
| Previous CIRU (v2.0.1) | 5/8 | 8/8 | 5/6 | 2/2 | 2/2 | 2/4 |
| CIRU v3 | 5/8 | 8/8 | 5/6 | 2/2 | 2/2 | 2/4 |
| Halo MTP3 | 6/8 | 8/8 | 5/6 | 2/2 | 2/2 | 3/4 |
| Profile | Hermes native full passes | Native mean points / 100 | Reviewed end states |
|---|---|---|---|
| Previous CIRU (v2.0.1) | 7/12 | 80.83 | 11/12 |
| CIRU v3 | 11/12 | 95.83 | 12/12 |
| Halo MTP3 | 11/12 | 95.83 | 12/12 |
V3 preserves the previous academic pass counts and improves Hermes native full passes from 7/12 to 11/12. Reviewed end states were 11/12, 12/12 and 12/12. Previous CIRU and Halo each have a memory-case wording artifact in the native grader; original grades are retained. All arms passed the short 10-task and long 8-task coding health checks, base and extended tests.
One request at a time. Native short tasks: temperature 0, seed 15035, nonthinking, 32,768-token output allowance. Hermes: temperature 0.6, top-p 0.95, top-k 20, thinking enabled and full remaining context. Long tasks return to one shared 63K history; seeding is charged separately. First trajectories only, no answer repair. Two Hermes rounds expose variation without establishing a failure probability.
Short-task native generation: previous CIRU 38.44, v3 39.09, Halo 43.97 tok/s. Long hard generation: 21.75, 35.93 and 31.25 tok/s respectively. Different generated lengths and tool actions affect elapsed time. Instrumented, interrupted and canceled captures are excluded.
Full report and case-level review ↓ · Structured measurements ↓
All four profiles passed 20/20 HumanEval base and extended tests on the separate canonical tasks 0–19 panel. Previous CIRU and v3 MTP6 produced identical output token IDs; v3 reduces their summed request time by 6.39%. This is a bounded nonthinking regression test, not a full 164-task result.
| Profile | Generated tokens | Prompt tok/s | Generation tok/s | Sum of request times · s |
|---|---|---|---|---|
| Previous CIRU | 3,179 | 148.53 | 53.33 | 75.49 |
| CIRU v3 MTP6 | 3,179 | 219.51 | 53.24 | 70.67 |
| CIRU v3 MTP2 | 3,212 | 226.25 | 39.63 | 91.57 |
| Halo MTP3 | 3,241 | 176.10 | 49.48 | 79.15 |
MTP6 retains high-acceptance coding speed; MTP2 lowers it to 39.63 tok/s. On the lower-acceptance fixed-output fixture, optional MTP2 instead measured 13.62 s at 4K and 180.87 s at 64K (29.39 / 24.88 generation tok/s). It remains an option, not the general default. The target model verifies the full vocabulary at either depth.
Recall checks recovered both keys with exact cached replay at approximately 8K and 64K for all three main arms. V3 passed 69 QSA mapping/state/guard cases, 33 ROCm operator reference cases, and 30 allocator tests with 198 assertions. A four-prefix diagnostic matched 15,892,480 F32 logits byte-for-byte; it does not establish universal equivalence or long-context task accuracy.
These are the completed H96, MTP depth-1 results behind the earlier release’s EvalScope scores. They used the same released weights, one request at a time and uncapped natural-EOS generation. They have not been rerun on v3.
| Dataset | Pass count | Coverage | Measured wall time |
|---|---|---|---|
| ARC-Challenge | 1143/1172 | Full dataset | 24m 42s |
| GPQA-Diamond | 46/50 | Sampled subset | 1h 33m 21s |
| MMLU-Pro | 61/70 | Sampled subset | 36m 37s |
| GSM8K | 97/100 | Sampled subset | 20m 19s |
| IFEval strict | 92/100 | Sampled subset | 22m 50s |
Quality stages together: 3h 17m 50s. Including the separate performance test: 3h 20m 46s. This is not the 25-minute v3 mixed panel above. No estimated v3 EvalScope wall time is presented as a measurement.
The following measurements retain their original runtime versions, hosts and protocols. They are historical v2 results, not new v3 measurements.
HumanEval cases 0–9 provide the coding prompts for this speed and draft-acceptance workload. Greedy generation, thinking disabled, one natural-EOS completion per case. All 30 requests completed with a 262,144-token server context for every package.
Runs were split between the two Ryzen AI MAX+ 395 hosts: CIRU on Ciru; Laurent and Unsloth sequentially on Sozo. Both used the performance CPU governor. All three use 256K server context. Model-specific settings are listed below.
Server decode time · higher is faster
Each point is one completed request
| Candidate / host | TG tok/s | Prompt tokens | Generated tokens | Total tokens | Request wall | Panel elapsed | MTP accepted |
|---|---|---|---|---|---|---|---|
| CIRU IU4ROCm10 · Ciru | 54.82 | 1,180 | 1,633 | 2,813 | 38.71 s | 39.81 s | 86.28% |
| Laurent FP4Vulkan · Sozo | 52.00 | 1,180 | 1,632 | 2,812 | 38.73 s | 39.79 s | 96.75% |
| Unsloth IQ4_XSVulkan · Sozo | 47.49 | 1,180 | 1,639 | 2,819 | 48.77 s | 49.61 s | 98.73% |
Total tokens = prompt + generated. Request wall includes prefill and streaming; panel elapsed also includes recorder overhead between requests and excludes model loading. This HumanEval panel measures speed and acceptance; it does not report a coding quality score.
| Task | CIRU IU4 TG | CIRU IU4 tokens | CIRU IU4 wall | Laurent FP4 TG | Laurent FP4 tokens | Laurent FP4 wall | Unsloth IQ4_XS TG | Unsloth IQ4_XS tokens | Unsloth IQ4_XS wall |
|---|---|---|---|---|---|---|---|---|---|
| HumanEval/0 | 48.32 | 173 | 4.84 s | 51.39 | 173 | 4.19 s | 47.71 | 176 | 5.17 s |
| HumanEval/1 | 51.91 | 232 | 5.36 s | 50.25 | 230 | 5.32 s | 45.59 | 239 | 6.69 s |
| HumanEval/2 | 54.71 | 97 | 2.52 s | 52.62 | 97 | 2.52 s | 48.73 | 97 | 3.35 s |
| HumanEval/3 | 53.47 | 157 | 3.96 s | 53.82 | 156 | 3.69 s | 46.95 | 157 | 4.85 s |
| HumanEval/4 | 54.26 | 169 | 4.23 s | 53.90 | 169 | 3.89 s | 47.98 | 169 | 4.99 s |
| HumanEval/5 | 58.11 | 154 | 3.36 s | 50.12 | 146 | 3.60 s | 47.53 | 144 | 4.39 s |
| HumanEval/6 | 53.17 | 214 | 4.87 s | 49.67 | 225 | 5.27 s | 47.41 | 221 | 6.12 s |
| HumanEval/7 | 57.89 | 112 | 2.65 s | 55.18 | 112 | 2.72 s | 48.67 | 112 | 3.67 s |
| HumanEval/8 | 59.16 | 163 | 3.61 s | 55.52 | 163 | 3.69 s | 48.63 | 163 | 4.81 s |
| HumanEval/9 | 63.63 | 162 | 3.29 s | 51.06 | 161 | 3.85 s | 47.62 | 161 | 4.73 s |
Four short prefixes span dialogue, mathematics, Python, and physics/engineering. The panel compares 64 distributions over 248,320 vocabulary entries, with perplexity calculated from 60 observed next-token losses.
Lower is closer to the reference
Lower is better · BF16 reference 2.0180
| Candidate | Mean KL ↓ | Median KL ↓ | P95 KL ↓ | Max KL ↓ | Top1 agreement ↑ | PPL ↓ |
|---|---|---|---|---|---|---|
| CIRU IU4 | 0.03045 | 0.00276 | 0.15453 | 0.33368 | 61/64 | 2.0840 |
| Laurent FP4 | 0.11223 | 0.00787 | 0.29765 | 1.91228 | 60/64 | 2.3838 |
| Unsloth IQ4_XS | 0.17457 | 0.00392 | 0.55722 | 3.84207 | 61/64 | 2.4293 |
This is a small numerical diagnostic. PPL is scoped to the 60-token slice. MTP is off; CIRU uses F16 KV / flash auto, and the retained competitor captures use Q8 KV / flash on. Two independent CIRU loads produced byte-identical logits.
MTP-off serving with seven exact source-prefix lengths and 128 generated tokens per request. All three sweeps ran on Sozo with a 262,144-token server context, preserving each package’s recorded settings.
PP · tok/s · higher is faster
TG · tok/s · MTP off
| Prompt tokens | CIRU IU4 PP | CIRU IU4 TG | Laurent FP4 PP | Laurent FP4 TG | Unsloth IQ4_XS PP | Unsloth IQ4_XS TG |
|---|---|---|---|---|---|---|
| 512 | 306.53 | 22.55 | 357.63 | 25.93 | 235.03 | 24.54 |
| 2,048 | 382.03 | 20.94 | 363.83 | 25.58 | 270.12 | 23.98 |
| 8,192 | 370.41 | 19.19 | 325.59 | 24.95 | 274.78 | 22.69 |
| 16,384 | 352.13 | 17.47 | 302.72 | 24.36 | 266.87 | 21.28 |
| 32,768 | 321.33 | 14.32 | 273.95 | 22.77 | 254.68 | 18.79 |
| 65,536 | 282.99 | 10.17 | 229.68 | 20.25 | 228.78 | 13.37 |
| 131,072 | 232.95 | 6.75 | 184.33 | 17.86 | 189.79 | 9.98 |
Prompt length on the horizontal axis
RAM is whole-system used memory. GTT is a shared-memory allocation counter; it overlaps system RAM and must not be added to it. Idle, peak and delta counters are retained in the JSON.
Cold request slots, temperature 0, seed 1234, EOS ignored for the 128-token output. One 512+32 warmup is excluded. Sweeps used the powersave CPU governor and measure a different workload from the performance-governor MTP panel above.
| Package | Runtime / backend | MTP recipe | Target KV | b / ub | Context | Threads |
|---|---|---|---|---|---|---|
| CIRU IU4 | CIRU RC2ROCm10 | Fixed depth 6 · p-min 0 · Q8 MTP head | F16 | 2048 / 512 | 262,144 | 8 |
| Laurent FP4 | Laurent 5e085d1Vulkan | Adaptive depth 2–4 · dedicated MTP head | Q8_0 | 2048 / 512 | 262,144 | 16 |
| Unsloth IQ4_XS | Unsloth d1a9235Vulkan | Depth 2 · shared Q8 MTP head | Runner default | 2048 / 512 | 262,144 | 16 |
Each canonical task is a single user message rendered by the model’s embedded chat template. Thinking is disabled with the native flag and enable_thinking=false; every rendered prompt ends with an empty, closed thinking block.
Temperature 0, top-k 1, top-p 1, min-p 0, neutral penalties, seed 123, one slot, no prompt reuse, no custom stops and no output-token cap. First samples are retained. Generated output contains no thinking blocks.
The measured packages include their weights, runtime, backend and live recommended serving settings. CIRU uses its fixed-six ROCm10 profile at 256K; Laurent uses adaptive 2–4 MTP; Unsloth uses its recommended shared-Q8 head at depth 2.
The context sweeps use their separately recorded target-only recipes. CIRU has F16 KV and 8 threads; Laurent has Q8 KV; Unsloth has Q8 KV, 16 threads, lazy mode off and explicit CPU placement of learned-memory weights. Both competitor cards leave batch/microbatch at their runner defaults, verified as 2048/512; CIRU explicitly uses 2048/512. Full commands are preserved.
Per-task TG is the server’s predicted_per_second. These runtimes exclude the first generated token from decode timing, so aggregate TG is Σ(generated tokens − 1) / Σ(decode seconds). Generated-token totals still retain every generated token. No arithmetic average of per-task rates is used.
CIRU’s panel was corrected to its released 256K context on 5 September; the earlier 16K diagnostic configuration is superseded. All 30 HumanEval rows were checked against the official performance SQLite store and raw SSE streams, including counts, natural EOS, MTP counters and non-thinking output. Both performance ledgers pass integrity verification. The 21 retained sweep rows have their own saved audit.
The HumanEval dataset SHA-256 is b796127e635a67f93fb35c04f4cb03cf06f38c8072ee7cee8833d7bee06979ef. The BF16 panel and reference hashes are included in BF16 data. CIRU’s retained binary hashes are checked before the run; no model weights were changed for this panel.
Packages: CIRU v2.0 · Agention / Laurent · Unsloth. Charts generated in Python with pyecharts and rendered with Apache ECharts.