How I got 2.2x more tokens per second from llama.cpp on Intel Arc
I spent a weekend benchmarking llama.cpp on my Xiaomi Book Pro 14. The machine has Intel's Core Ultra X7 358H and Arc B390 integrated graphics, with 30 GiB of unified memory. That last part matters. It means the GPU can address most of system RAM, which changes what you can load.
I was running Qwen3.6-35B, a 35B parameter mixture-of-experts model with about 3B active parameters, quantized to Q4_K_M. The weights are 21.1 GiB. Add 32k context with q4_0 KV cache and you are looking at roughly 26 GiB total, leaving about 5 GiB free on my 30 GiB machine.
The stock configuration had CPU_MOE=1, which forces all MoE expert layers onto the CPU. I assumed this was necessary because the Arc B390 only exposes 23.7 GiB of device memory. It turns out the full model fits fine at 32k context. Running those experts on the GPU instead is worth a staggering amount of speed.
The single biggest setting: CPU_MOE
I ran llama-bench with the MoE CPU offload parameter set to different values. Everything else stayed the same: 8 threads, batch 512, flash attention on, q4_0 KV cache.
| MoE layers on CPU | Prompt processing (t/s) | Token generation (t/s) |
|---|---|---|
| 0 (full GPU) | 686 | 32.5 |
| 8 | 395 | 23.4 |
| 24 | 330 | 20.3 |
| 41 (all CPU) | 293 | 14.7 |
Full GPU offload is 2.2x faster at generation and 2.3x faster at prompt processing compared to CPU_MOE=1. That is not a marginal improvement. It changes the usability of the model entirely.
I confirmed this on the live server. A 400-token completion jumped from 15-17 t/s to 33-36 t/s. The old README warned that CPU_MOE was required, citing a 10% speed cost. The real cost is 55% of your decode throughput.
The catch is memory. Full offload leaves about 5 GiB free. It works, but it is tight. If you try to load a second model or push context past 64k, you will run out. I added a guard in my launch script that switches back to CPU_MOE=1 when context exceeds 49k tokens. At that point the KV cache alone would push past 8 GiB and there would not be enough room for the weights.
Speculative decoding: n=2 beats n=5
I had been running MTP speculative decoding with spec-draft-n-max=5. The idea is that the model drafts five tokens at once, verifies them in parallel, and accepts the correct ones. More drafts should mean more accepted tokens per verification step, right?
Not on this hardware. Draft cost scales with n, and at n=5 the extra work eats most of the gain.
| Speculation depth | Tokens per second | Gain over no speculation |
|---|---|---|
| 0 (none) | 29.0 | β |
| 1 | 36.9 | +27% |
| 2 | 38.3 | +32% |
| 3 | 36.8 | +27% |
| 5 | 31.5 | +9% |
n=2 is the sweet spot. It gives 32% more throughput for free, and the draft model is small enough to load without memory issues. n=5 actually failed to allocate a 236 MB compute buffer under full GPU offload, forcing me to keep one layer on CPU just to make it fit.
The acceptance rate at n=5 was around 0.4, meaning about 40% of drafted tokens were accepted. At n=2 and n=3 the acceptance rate was higher and net speed was better because the draft step itself is cheap.
Batch size for long prompts
The launch script had an auto-clamp that set batch and ubatch to 512 whenever context exceeded 12k tokens. This was presumably a memory safety measure, but on my machine it left performance on the table for long prompts.
| Batch size | pp512 | pp4096 |
|---|---|---|
| 512 | 688 | 523 |
| 1024 | 683 | 601 |
| 2048 | 684 | 639 |
| 4096 | β | OOM |
Short prompts are unaffected. At 4096 tokens, batch 2048 is 22% faster than batch 512. On the live server with a real 30k-token prompt, the difference was 130 t/s versus 144 t/s. Batch 4096 crashed with an out-of-memory error, so 2048 is the ceiling.
I changed the auto-clamp from 512 to 2048 in the launch script.
Things I tried that did not matter
Thread count. I tested 2, 6, 8, 10, 12, and 16 threads. Prompt processing varied within 1%. Token generation stayed at 32.4-32.7 t/s across every thread count. The workload is GPU-bound. CPU threading is irrelevant here.
CPU governor. I switched between performance and powersave on the intel_pstate driver. The numbers were identical, both in the GPU-bound case and when I forced everything to the CPU. The powersave governor on intel_pstate does not actually throttle like the old acpi-cpufreq driver. The scaling governor label is misleading. My launch script used to warn about this. On this machine, it does not matter.
KV cache quantization. q4_0, q8_0, and f16 all performed within about 1% of each other for both prompt processing and token generation. The only difference is memory usage. q4_0 uses roughly half the KV memory of f16, which at 32k context is the difference between fitting and not fitting. Keep q4_0.
--no-kv-offload. This is catastrophic. 133 t/s prompt processing versus 691. 2.85 t/s generation versus 32.7. On unified memory hardware, keeping the KV cache on the CPU means every cache hit crosses the memory bus instead of staying in VRAM. Do not use this.
--op-offload and -ngl auto. Both were within noise of the defaults. The auto-detection already offloads everything it can.
The memory budget
The reason full offload works on 23.7 GiB of device memory:
Model weights: 21.10 GiB
KV cache @ 32k, q4_0: ~4.2 GiB
Compute buffers and server: ~1 GiB
ββββββββββββββββββββββββββββββββββββββββ
Total: ~26.4 GiBThat leaves about 3.6 GiB below the device heap limit and roughly 5 GiB of system RAM free. I measured 24-25 GiB used and 5.2-6.3 GiB available on the live server. zram swap barely activated. The margin is small but real.
At 64k context, the KV cache jumps to about 8.4 GiB, pushing total memory past 30 GiB. That is where CPU_MOE=1 becomes necessary again. Not because the GPU cannot handle the model, but because there is not enough room for everything at once.
What I changed
Three settings, all in the launch script and model config:
CPU_MOEdefault from 1 to 0. This is the big one.- Batch and ubatch defaults from 512 to 2048. The auto-clamp now caps at 2048 instead of 512.
spec-draft-n-maxfrom 5 to 2 in the model config.
After applying these and restarting the server with no environment overrides, the launch line carries --batch-size 2048 --ubatch-size 2048 and no --cpu-moe. The bench model hits 38.7 t/s, up from 15-17. The 30k-token prompt processes at 145 t/s, up from 130. Minimum available memory stays above 4 GiB with no out-of-memory errors.
Reproduction
If you want to run the same benchmarks:
# MoE offload matrix
./build/bin/llama-bench -m models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf \
-t 8 -p 512 -n 128 -r 3 -b 512 -ub 512 -fa on -ctk q4_0 -ctv q4_0 -ngl 99 -ncmoe N
# Batch size sweep
./build/bin/llama-bench -m models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf \
-t 8 -p 4096 -n 1 -r 2 -fa on -ctk q4_0 -ctv q4_0 -ngl 99 -b 2048 -ub 2048 -ncmoe N
# Full GPU server (the fast config)
CPU_MOE=0 CTX=32768 PORT=8081 ./llama-server-start-multi.sh \
-- --batch-size 2048 --ubatch-size 2048The build was llama.cpp 0.5.0-dev (build 158), compiled with -DGGML_VULKAN=ON -DGGML_NATIVE=ON on GNU 16.2.1. Your mileage will vary on different hardware. The principle holds: on unified memory systems, offload everything you can to the GPU, keep batch sizes reasonable, and do not trust defaults that were tuned for discrete GPUs with limited VRAM.