Summary · Technical report
Local inference · Hardware reportVersion of 30 April 2026 · data to 27 April
Gemma 4 · Apple Silicon · integrating Claude Code and Copilot CLI
The Real Bottleneck Is KV Cache and Prefill, Not the Context Limit
This report works out what hardware is needed to run Gemma 4 locally behind an agentic CLI tool. The conclusion first: in an agentic workflow what determines how it feels is prefill speed and KV cache footprint, not how much context the model can accept. That is also why the 26B-A4B mixture-of-experts model, with only 4B parameters active, is the sweet spot for local agentic work.
- 60–70%KV cache saved by Gemma's 5:1 attention
- ÷2throughput roughly halves for each further 25% offloaded to CPU
- ~99 slatency of one tool-call cycle on an M4 Pro with 24 GB
- 5–7×how much more a comparable PC costs to run around the clock
About these figures
Measurements of tokens per second, KV cache and power draw come from community sources: llama.cpp, Unsloth, LLMCheck, Hardware Corner and r/LocalLLaMA. Combinations that were not measured are marked as estimates. None of this is purchasing advice, and prices and models change over time.
Context · Models and memory · Content 1 / 4
First, a Correction: Gemma 4 Is Not Named 1B/4B/12B/27B a common mix-up
The Gemma 4 released by Google DeepMind on 2 April 2026 is named E2B, E4B, 26B-A4B and 31B. The 1B, 4B, 12B and 27B sizes belong to Gemma 3 (2025 into early 2026), which is still available from Ollama. The tables below cover both generations.
| Commonly said | Gemma 4 SKU | Architecture | Max context |
|---|---|---|---|
| 1B | E2B | Dense with PLE | 128K |
| 4B | E4B | Dense with PLE | 128K |
| 12B | No equivalent; nearest is 26B-A4B | Mixture of experts — 128 experts, 8 active | 256K |
| 27B | 31B dense, the flagship | Dense | 256K |
PLE is per-layer embeddings. E2B and E4B perform like dense models but carry more parameters than the name suggests; in effective terms they are about 2.3B and 4.5B.
GGUF Weight Sizes Unsloth Dynamic 2.0
| Model | Q4_K_M | Q8_0 | BF16 |
|---|---|---|---|
| Gemma 4 E2B (~2.3B effective) | ~1.6 GB | ~2.6 GB | ~4.6 GB |
| Gemma 4 E4B (~4.5B effective) | ~3.0 | ~5.0 | ~9.0 |
| Gemma 4 26B-A4B MoE | 16.9 | 26.9 | 50.5 |
| Gemma 4 31B dense | 18.3 | 32.6 | 61.4 |
| Gemma 3 12B | 6.6 | ~13 | 24 |
| Gemma 3 27B | 14.1–15.1 | ~28 | 54 |
KV Cache: It Can Be Cut to a Third Gemma's 5:1 local-to-global attention
KV size is 2 × L × H_kv × T × d_head × bytes_per_element, where bytes per element is 2 for FP16, 1 for Q8_0 and 0.5 for Q4_0. In practice, enabling --cache-type-k q8_0 -fa on in llama.cpp, or OLLAMA_KV_CACHE_TYPE=q8_0 in Ollama, takes the KV cache for a 32K context from roughly 15 GB down to about 5 GB.
- 60–70%The saving from Gemma 3 and 4's 5:1 local-to-global attention, with a 1024 sliding window, against pure global attention.
- ×0.5Q8_0 KV quantisation halves it outright.
- ~30%Flash Attention shrinks the activation buffer, and Gemma 4's hybrid local/global attention depends on its sliding window.
Unified Memory Needed on macOS weights plus KV at FP16 with Flash Attention, plus 4–6 GB for the OS
| Quantisation (weights) | 8K | 32K | 128K | 256K |
|---|---|---|---|---|
| 26B-A4B Q4 (16.9 GB) | 24 | 24 | 32 | 32 |
| 26B-A4B Q8 (26.9 GB) | 32 | 48 | 48 | 64 |
| 31B Q4 (18.3 GB) | 24 | 32 | 32 | 48 |
| 31B Q8 (32.6 GB) | 48 | 48 | 64 | 96 |
| 31B BF16 (61.4 GB) | 96 | 96 | 128 | 192 |
Hardware · And the traps · Content 2 / 4
Three Known Traps on Apple Silicon Metal and Flash Attention
It simply hangs
Ollama with Gemma 4 and Flash Attention
- Anything over about 500 tokens of prompt hangs the whole thing.
- Codex CLI's system prompt is around 27K tokens, so it almost always triggers this.
- The fix is to connect to llama.cpp directly with
-fa on -ctk q8_0 -ctv q8_0.
Bandwidth went backwards
The M3 Pro trap
- Apple cut M3 Pro memory bandwidth from 200 GB/s to 150 GB/s.
- For some users that made it slower than an M2 Pro on LLM work.
A bonus
The MLX backend
- On models below 14B it is 20–87% faster than llama.cpp.
- Ollama 0.19 and later enable it automatically on Macs with 32 GB or more, improving decode by around 93% for most models.
For Comparison: NVIDIA VRAM and Speed Q4_K_M · 4K context · tokens per second
| Card | 8B | 14B | 26B MoE | 31B dense |
|---|---|---|---|---|
| RTX 3060 12 GB | 42 | 23–29 | OOM | OOM |
| RTX 4060 Ti 16 GB | 50 | 35 | 60–70 | OOM |
| RTX 4080 16 GB | 75 | 50 | 90–110 | OOM |
| RTX 4090 24 GB | 104 | 75–95 | 140–150 | 7.8 * |
| RTX 5090 32 GB | 130–150 | 100–120 | 180+ | 35 |
* Running 31B Q4 on a 4090 forces part of the model into system RAM, and a workstation with multi-channel DDR5 and a 64-core CPU beats it at 8.8 tok/s. That single cell is the most important in the table: it shows that not fitting is not a small slowdown but a change of order.
What Insufficient VRAM Costs the rule for hybrid inference
Take Qwen 3 8B Q4 on an RTX 4060 with 8 GB: offloading 25 of 37 layers drops it from 40.58 to 8.62 tokens per second — 4.7 times slower — and all it buys is VRAM falling from 7.2 GB to 4.8 GB. As a rule, each additional 25% offloaded to CPU roughly halves throughput, and an agentic workflow with many rounds of tool calling multiplies that gap.
- 30 seconds becomes 2–3 minutesWhat a 30-second tool-call cycle actually feels like when VRAM is short.
- ~4 tok/sSpeed with num_gpu=0, CPU only — effectively unusable in an agentic loop.
Pivot · CLI integration and latency · Content 3 / 4
Claude Code CLI and Copilot CLI both take a local model, over different protocols
| Item | Claude Code | Copilot CLI |
|---|---|---|
| Native Ollama support | Yes, from 0.14 in January 2026 | Yes, bring-your-own-key, April 2026 |
| Protocol | Anthropic Messages | OpenAI Chat Completions |
| Minimum context suggested | 64K | 128K |
| Tool calling | Required | Required, with streaming |
| Offline mode | Partial | Full, with COPILOT_OFFLINE=true |
| Suggested local models | GLM-4.7-flash, Gemma 4 26B-A4B | qwen3-coder, glm-5 |
How Local Models Score at Agentic Work tool-calling success and multi-step ability
| Model | Size | Tool calling | Multi-step |
|---|---|---|---|
| GLM-4.7-flash | 30B MoE | 98% | Strong |
| Qwen3-Coder 30B-A3B | 30B MoE | 93% | Very strong (SWE 71.3) |
| Devstral-2 24B | 24B | 88% | Moderate to strong |
| Gemma 4 31B Dense | 31B | 85% | Strong (τ2 86.4) |
| Gemma 4 26B-A4B | 26B MoE | 85% * | Moderate to strong |
| Qwen3 14B | 14B | 75% | Moderate |
| Llama 3.1 8B | 8B | 50% | Weak |
Known Tool-Calling Problems in Ollama each one breaks the whole agentic flow
- Tool calling with Qwen 3.5 35B-A3B is entirely broken on Ollama — the renderer and parser misalign on
<think>tags and tool-call prefixes, andrepeat_penaltyandpresence_penaltyare silently ignored (ollama/ollama#14493). Connect to llama.cpp directly with--jinjainstead. - Streaming tool calls with Gemma 4: v0.20.3's streaming path puts
tool_callsinto the reasoning channel, so Codex and Claude Code fail to parse them entirely. v0.20.5 or later is needed. - The default context is far too small: Ollama defaults to
num_ctx = 2048, nowhere near enough for Claude Code, whose system prompt is around 27K. SetOLLAMA_CONTEXT_LENGTH=131072or specify it in the Modelfile. - CUDA 13.2 has quality problems with Gemma 4 GGUF; Unsloth recommends dropping to 12.8.
What One Tool-Call Cycle Actually Costs this is where the feel comes from
| Setup | Latency | Conditions |
|---|---|---|
| Mac M4 Pro, 24 GB | ~99 s | llama.cpp with Gemma 4 26B Q4, including 27K tokens of prefill |
| DGX Spark, 128 GB | ~52 s | Blackwell GB10 with Ollama and Gemma 4 31B Q4, via Codex CLI |
| Cloud Sonnet 4 | ~12 s | at high reasoning; Opus 4.5 is around 15 s |
Adjust the timeout
Claude Code's default stream_idle_timeout_ms of 600,000 — ten minutes — is only barely enough for local inference. In practice raise it to 1,800,000, thirty minutes, to cover prefill on a Mac. One further piece of community advice: pin your llama.cpp version, since a 3.3× speed regression has appeared between builds.
Resolution · What to buy, and what it costs to keep · Content 4 / 4
Mac mini Configurations assuming it runs continuously
| Tier | Configuration | Taiwan price | Runs |
|---|---|---|---|
| Entry | M4, 16 GB / 256 GB | NT$19,900 | An 8B agent |
| Entry plus | M4, 16 GB / 512 GB | NT$26,900 | 8B, with room for several models |
| Mid ★ | M4, 24 GB / 512 GB | NT$33,900 | 14B Q4 with 32K context |
| Higher | M4 Pro 12-core, 24 GB / 512 GB, 273 GB/s | NT$46,900 | 26B-A4B Q4 |
| Higher plus | M4 Pro 14-core / 20-core GPU | NT$53,900 | 26B-A4B Q4 and above |
| Ideal dev box | M4 Pro 14-core / 20-core GPU, 48 GB / 1 TB | NT$66,900 | 30B at Q5 or Q8 — over budget |
As of 30 April 2026 the Mac mini is still on the M4 and M4 Pro released in October 2024. Education pricing takes off roughly NT$2,000–4,500 depending on model. On laptops: the M5 Max does prompt processing about 2.3 times faster than the M4 Max and decodes tokens about 28% faster, and its new neural accelerators make a marked difference to time-to-first-token — which for agentic use is exactly where the bottleneck sits.
Running Costs and Three-Year Total half inference, half idle · NT$3.5 per kWh
| Machine | Idle W | Inference W | Monthly power | Three-year total |
|---|---|---|---|---|
| Mac mini M4, 16 GB | 3–4 | 35–45 | NT$60 | NT$22,000 |
| Mac mini M4 Pro, 24 GB | 25–30 | 78–92 | NT$140 | NT$52,000 |
| PC with RTX 4060 Ti 16 GB | ~80 | ~280 | ~NT$400 | NT$52,400 |
| PC with RTX 4090 24 GB | 80–110 | 380–450 | ~NT$750 | NT$110,000 |
| PC with RTX 5090 32 GB | ~110 | ~500 | ~NT$900 | NT$130,000+ |
For an assistant that runs continuously
For the same money a Windows machine can win on raw specification, but costs five to seven times as much to run around the clock. Between NT$20,000 and NT$50,000 the Mac mini has almost no competition: quiet, low-draw, and markedly cheaper over three years than a PC at the same price. The limitation to note is that an external GPU is not an option — the only way to more usable memory is a model with more of it, and the choice is fixed at purchase.
The middle option worth recommending
The M4 with 24 GB and 512 GB, at NT$33,900, runs 14B Q4 with 32K context and an agent entirely on the GPU, for about NT$2,200 of electricity over three years. If the budget stretches to the NT$46,900 M4 Pro 12-core, its 273 GB/s of bandwidth takes 14B decode from 25 to 50 tokens per second — the point at which an agent goes from usable to smooth, and the largest single step in perceived difference anywhere on the curve.