Ciru Inference Labllm.ciru.ai / research
Crown Citadel Research ReportV3 RESULTS · 09 SEP 2026

Qwen3.8 Flash.
V3. Measured end to end.

CIRU v3 on Strix Halo: faster long-context serving, measured task completion times, and agent-quality checks. Compare the previous CIRU runner and Halo, then explore the retained v2 evidence.

CIRU v3ROCm10 · MTP6Previous CIRUv2.0.1 · MTP6HaloVulkan · MTP3
V3 · 64K request time23.9% lessvs previous CIRU · 239.99 → 182.57 s
V3 · mixed panel25m 18s13.6% less time than previous CIRU
V3 · coding regression20 / 20HumanEval base and extended tests
V3 · capacity verified261,888real input tokens + 128 generated
01 / V3 serving · 09 SEP 2026

Faster long-context requests

Serving CSV ↓

Same CIRU weights, updated runtime and serving profile. At 64K, v3 reduces whole-request time by 23.93%, raises prompt throughput by 29.99%, and raises generation throughput by 81.66% versus the previous CIRU runner. At 4K, whole-request time falls 11.80%.

All three arms ran sequentially on Ciru: Ryzen AI Max+ 395 / gfx1151, 128 GB unified memory, NixOS. Identical input token IDs, cold prompt cache, 128 output tokens, one slot and 262,144 context capacity. Each load has an excluded warmup. Previous CIRU is the locally qualified v2.0.1 runner; it was not separately published as a tag.

Fixed-output serving experiment · arithmetic mean latency; pooled native PP and TG
Input tokensProfilePrompt tok/sGeneration tok/sFirst piece · sWhole request · sRequests
4,096Previous CIRU392.0022.5210.7016.342
4,096CIRU v3455.6524.609.2514.416
4,096Halo381.4935.3011.0914.694
65,536Previous CIRU284.4913.33230.46239.991
65,536CIRU v3369.8124.22177.32182.573
65,536Halo263.4223.28248.91254.372

64K whole-request time

Seconds · lower is faster

Complete mixed benchmark

Minutes · includes tools and scoring

Previous CIRU uses MTP6, b2048/u512; v3 uses MTP6, b1024/u1024. Both use F16 target KV and Q8 draft KV. Halo uses its unmodified Vulkan runner, Unsloth UD-Q4_K_XL target, EasiiX Q8 head and screened MTP3 profile. Halo generates faster at 4K; the small 4K whole-request difference is not a robust general v3 win. Three v3 loads, two Halo loads and one previous-CIRU load do not establish confidence intervals. This measures complete serving profiles, not runtime changes in isolation.

Sampler, memory and full-capacity result

Nonthinking, temperature 0.7, top-p 0.8, top-k 20, min-p 0, presence penalty 1.5, repeat penalty 1, frequency penalty 0, seed 123. EOS is honored; all reported requests produced 128 tokens. Native TG excludes the first output token; first-piece latency is the first streamed content-field event.

Sampled telemetry · RAM, GTT and VRAM overlap and must not be added
InputProfilePeak system RAM · GiBMax RAM increase · GiBPeak GTT · GiBPeak VRAM · GiB
4,096Previous CIRU103.432.1588.580.52
65,536Previous CIRU101.320.7288.580.52
4,096CIRU v3106.022.4390.290.53
65,536CIRU v3103.660.7490.290.52
4,096Halo104.882.7793.081.85
65,536Halo103.720.6093.171.99

The v3 capacity check completed 261,888 input + 128 output tokens: 257.44 prompt tok/s, 18.00 generation tok/s, 1017.38 s to first piece and 1024.44 s whole request. One CIRU-only request establishes serving capacity, not full-context answer accuracy or a 256K comparison with Halo.

02 / Task quality and elapsed time

Mixed tasks, agents and long history

Wall times CSV ↓

The additional Ornith difficulty panel combines 22 short IFEval/GSM8K/HumanEval tasks, six Hermes scenarios repeated twice, eight hard tasks with a shared 63K-token history, and short/long coding health checks. It uses cases selected from earlier disagreements and failures; scores do not estimate general benchmark accuracy.

Actual elapsed time, including applicable scorer/tool overhead. Stage rows are components of the complete panel; coding health and other overhead also contribute.
StagePrevious CIRUCIRU v3Halo MTP3
Complete panel, after readiness29m 17s25m 18s24m 34s
Model load, additional0m 31s0m 31s0m 34s
Short scored stage5m 20s5m 21s4m 36s
Hermes, two rounds combined12m 28s11m 38s10m 08s
63K history seeding3m 26s2m 43s3m 21s
Long hard stage, history already loaded4m 44s2m 42s3m 23s
All generated tokens33,29433,45131,058

V3 takes 13.59% less total time than previous CIRU, while Halo finishes 43.89 seconds sooner than v3. Long hard-stage time falls 42.98% versus previous CIRU, but output lengths differ. Short-stage time is effectively unchanged. These are complete workload timings, not equal-output generation speedups.

Native first-sample scores · short and long tasks are different subsets
ProfileShort IFEvalShort GSM8KShort HumanEvalLong IFEvalLong GSM8KLong HumanEval
Previous CIRU (v2.0.1)5/88/85/62/22/22/4
CIRU v35/88/85/62/22/22/4
Halo MTP36/88/85/62/22/23/4
Six scenarios × two repetitions; reviewed end states do not replace native scores
ProfileHermes native full passesNative mean points / 100Reviewed end states
Previous CIRU (v2.0.1)7/1280.8311/12
CIRU v311/1295.8312/12
Halo MTP311/1295.8312/12

V3 preserves the previous academic pass counts and improves Hermes native full passes from 7/12 to 11/12. Reviewed end states were 11/12, 12/12 and 12/12. Previous CIRU and Halo each have a memory-case wording artifact in the native grader; original grades are retained. All arms passed the short 10-task and long 8-task coding health checks, base and extended tests.

Hard-panel protocol and review details

One request at a time. Native short tasks: temperature 0, seed 15035, nonthinking, 32,768-token output allowance. Hermes: temperature 0.6, top-p 0.95, top-k 20, thinking enabled and full remaining context. Long tasks return to one shared 63K history; seeding is charged separately. First trajectories only, no answer repair. Two Hermes rounds expose variation without establishing a failure probability.

Short-task native generation: previous CIRU 38.44, v3 39.09, Halo 43.97 tok/s. Long hard generation: 21.75, 35.93 and 31.25 tok/s respectively. Different generated lengths and tool actions affect elapsed time. Instrumented, interrupted and canceled captures are excluded.

Full report and case-level review ↓ · Structured measurements ↓

03 / MTP and regression checks

Draft depth depends on the workload

All four profiles passed 20/20 HumanEval base and extended tests on the separate canonical tasks 0–19 panel. Previous CIRU and v3 MTP6 produced identical output token IDs; v3 reduces their summed request time by 6.39%. This is a bounded nonthinking regression test, not a full 164-task result.

One sample per task, EvalPlus v0.1.10, 4096-token cap; truncations fail. Summed request time excludes grading.
ProfileGenerated tokensPrompt tok/sGeneration tok/sSum of request times · s
Previous CIRU3,179148.5353.3375.49
CIRU v3 MTP63,179219.5153.2470.67
CIRU v3 MTP23,212226.2539.6391.57
Halo MTP33,241176.1049.4879.15

MTP6 retains high-acceptance coding speed; MTP2 lowers it to 39.63 tok/s. On the lower-acceptance fixed-output fixture, optional MTP2 instead measured 13.62 s at 4K and 180.87 s at 64K (29.39 / 24.88 generation tok/s). It remains an option, not the general default. The target model verifies the full vocabulary at either depth.

Recall checks recovered both keys with exact cached replay at approximately 8K and 64K for all three main arms. V3 passed 69 QSA mapping/state/guard cases, 33 ROCm operator reference cases, and 30 allocator tests with 198 assertions. A four-prefix diagnostic matched 15,892,480 F32 logits byte-for-byte; it does not establish universal equivalence or long-context task accuracy.

04 / Historical quality · 29 AUG 2026

EvalScope baseline: 3h 17m 50s

These are the completed H96, MTP depth-1 results behind the earlier release’s EvalScope scores. They used the same released weights, one request at a time and uncapped natural-EOS generation. They have not been rerun on v3.

1,492 items · actual per-stage start/end timestamps, excluding setup and failed attempts
DatasetPass countCoverageMeasured wall time
ARC-Challenge1143/1172Full dataset24m 42s
GPQA-Diamond46/50Sampled subset1h 33m 21s
MMLU-Pro61/70Sampled subset36m 37s
GSM8K97/100Sampled subset20m 19s
IFEval strict92/100Sampled subset22m 50s

Quality stages together: 3h 17m 50s. Including the separate performance test: 3h 20m 46s. This is not the 25-minute v3 mixed panel above. No estimated v3 EvalScope wall time is presented as a measurement.

Historical v2.0 comparison · CIRU, Laurent & Unsloth · 05 SEP 2026

The following measurements retain their original runtime versions, hosts and protocols. They are historical v2 results, not new v3 measurements.

CIRU · non-thinking MTP54.82 tok/sHumanEval 0–9 · aggregate TG
CIRU · BF16 mean KL ↓0.0304564 full-vocabulary distributions
CIRU · panel PPL ↓2.084060 observed next tokens
CIRU · 128K prefill232.95 tok/sMTP off · exact source tokens
01 / Served generation

Non-thinking MTP speed

Per-task CSV ↓

HumanEval cases 0–9 provide the coding prompts for this speed and draft-acceptance workload. Greedy generation, thinking disabled, one natural-EOS completion per case. All 30 requests completed with a 262,144-token server context for every package.

CIRU: 54.82 tok/s+5.4% versus Laurent and +15.5% versus Unsloth in these runs.

Runs were split between the two Ryzen AI MAX+ 395 hosts: CIRU on Ciru; Laurent and Unsloth sequentially on Sozo. Both used the performance CPU governor. All three use 256K server context. Model-specific settings are listed below.

Aggregate generation rate

Server decode time · higher is faster

HumanEval 0–9

Each point is one completed request

Ten completed requests per candidate. TG uses server decode time; wall time includes prefill.
Candidate / hostTG tok/sPrompt tokensGenerated tokensTotal tokensRequest wallPanel elapsedMTP accepted
CIRU IU4ROCm10 · Ciru54.821,1801,6332,81338.71 s39.81 s86.28%
Laurent FP4Vulkan · Sozo52.001,1801,6322,81238.73 s39.79 s96.75%
Unsloth IQ4_XSVulkan · Sozo47.491,1801,6392,81948.77 s49.61 s98.73%

Total tokens = prompt + generated. Request wall includes prefill and streaming; panel elapsed also includes recorder overhead between requests and excludes model loading. This HumanEval panel measures speed and acceptance; it does not report a coding quality score.

All ten tasks: TG, generated tokens and wall time
Per-task measurements · TG in tok/s · tokens are generated tokens
TaskCIRU IU4 TGCIRU IU4 tokensCIRU IU4 wallLaurent FP4 TGLaurent FP4 tokensLaurent FP4 wallUnsloth IQ4_XS TGUnsloth IQ4_XS tokensUnsloth IQ4_XS wall
HumanEval/048.321734.84 s51.391734.19 s47.711765.17 s
HumanEval/151.912325.36 s50.252305.32 s45.592396.69 s
HumanEval/254.71972.52 s52.62972.52 s48.73973.35 s
HumanEval/353.471573.96 s53.821563.69 s46.951574.85 s
HumanEval/454.261694.23 s53.901693.89 s47.981694.99 s
HumanEval/558.111543.36 s50.121463.60 s47.531444.39 s
HumanEval/653.172144.87 s49.672255.27 s47.412216.12 s
HumanEval/757.891122.65 s55.181122.72 s48.671123.67 s
HumanEval/859.161633.61 s55.521633.69 s48.631634.81 s
HumanEval/963.631623.29 s51.061613.85 s47.621614.73 s
02 / Numerical fidelity

How close are the logits to BF16?

BF16 CSV ↓

Four short prefixes span dialogue, mathematics, Python, and physics/engineering. The panel compares 64 distributions over 248,320 vocabulary entries, with perplexity calculated from 60 observed next-token losses.

Forward KL from BF16

Lower is closer to the reference

Perplexity on the panel

Lower is better · BF16 reference 2.0180

BF16 reference PPL: 2.0180 · 64 distributions / 60 observed next tokens
CandidateMean KL ↓Median KL ↓P95 KL ↓Max KL ↓Top1 agreement ↑PPL ↓
CIRU IU40.030450.002760.154530.3336861/642.0840
Laurent FP40.112230.007870.297651.9122860/642.3838
Unsloth IQ4_XS0.174570.003920.557223.8420761/642.4293

This is a small numerical diagnostic. PPL is scoped to the 60-token slice. MTP is off; CIRU uses F16 KV / flash auto, and the retained competitor captures use Q8 KV / flash on. Two independent CIRU loads produced byte-identical logits.

03 / Context scaling

512 tokens to 128K

Sweep CSV ↓

MTP-off serving with seven exact source-prefix lengths and 128 generated tokens per request. All three sweeps ran on Sozo with a 262,144-token server context, preserving each package’s recorded settings.

Prompt processing

PP · tok/s · higher is faster

Generation after prefill

TG · tok/s · MTP off

MTP off · 128 generated tokens per request · PP and TG in tok/s
Prompt tokensCIRU IU4 PPCIRU IU4 TGLaurent FP4 PPLaurent FP4 TGUnsloth IQ4_XS PPUnsloth IQ4_XS TG
512306.5322.55357.6325.93235.0324.54
2,048382.0320.94363.8325.58270.1223.98
8,192370.4119.19325.5924.95274.7822.69
16,384352.1317.47302.7224.36266.8721.28
32,768321.3314.32273.9522.77254.6818.79
65,536282.9910.17229.6820.25228.7813.37
131,072232.956.75184.3317.86189.799.98
Explore request latency and memory

Latency & memory

Prompt length on the horizontal axis

RAM is whole-system used memory. GTT is a shared-memory allocation counter; it overlaps system RAM and must not be added to it. Idle, peak and delta counters are retained in the JSON.

Cold request slots, temperature 0, seed 1234, EOS ignored for the 128-token output. One 512+32 warmup is excluded. Sweeps used the powersave CPU governor and measure a different workload from the performance-governor MTP panel above.

04 / Reproduction

Settings, sources and records

All results JSON ↓
HumanEval serving profiles. Full commands, binary hashes and templates are in the protocol download.
PackageRuntime / backendMTP recipeTarget KVb / ubContextThreads
CIRU IU4CIRU RC2ROCm10Fixed depth 6 · p-min 0 · Q8 MTP headF162048 / 512262,1448
Laurent FP4Laurent 5e085d1VulkanAdaptive depth 2–4 · dedicated MTP headQ8_02048 / 512262,14416
Unsloth IQ4_XSUnsloth d1a9235VulkanDepth 2 · shared Q8 MTP headRunner default2048 / 512262,14416

The HumanEval request

Each canonical task is a single user message rendered by the model’s embedded chat template. Thinking is disabled with the native flag and enable_thinking=false; every rendered prompt ends with an empty, closed thinking block.

Temperature 0, top-k 1, top-p 1, min-p 0, neutral penalties, seed 123, one slot, no prompt reuse, no custom stops and no output-token cap. First samples are retained. Generated output contains no thinking blocks.

What is being compared

The measured packages include their weights, runtime, backend and live recommended serving settings. CIRU uses its fixed-six ROCm10 profile at 256K; Laurent uses adaptive 2–4 MTP; Unsloth uses its recommended shared-Q8 head at depth 2.

The context sweeps use their separately recorded target-only recipes. CIRU has F16 KV and 8 threads; Laurent has Q8 KV; Unsloth has Q8 KV, 16 threads, lazy mode off and explicit CPU placement of learned-memory weights. Both competitor cards leave batch/microbatch at their runner defaults, verified as 2048/512; CIRU explicitly uses 2048/512. Full commands are preserved.

Timing conventions and audit

Per-task TG is the server’s predicted_per_second. These runtimes exclude the first generated token from decode timing, so aggregate TG is Σ(generated tokens − 1) / Σ(decode seconds). Generated-token totals still retain every generated token. No arithmetic average of per-task rates is used.

CIRU’s panel was corrected to its released 256K context on 5 September; the earlier 16K diagnostic configuration is superseded. All 30 HumanEval rows were checked against the official performance SQLite store and raw SSE streams, including counts, natural EOS, MTP counters and non-thinking output. Both performance ledgers pass integrity verification. The 21 retained sweep rows have their own saved audit.

The HumanEval dataset SHA-256 is b796127e635a67f93fb35c04f4cb03cf06f38c8072ee7cee8833d7bee06979ef. The BF16 panel and reference hashes are included in BF16 data. CIRU’s retained binary hashes are checked before the run; no model weights were changed for this panel.

Packages: CIRU v2.0 · Agention / Laurent · Unsloth. Charts generated in Python with pyecharts and rendered with Apache ECharts.