Hack Day Starter
Technical WhitepaperHardware Engineering

Ollama Hardware Requirements Guide

Running large language models locally is memory-bound rather than compute-bound. This guide covers how memory architectures, VRAM boundaries, and KV cache expansion dictate real-world performance.

1. The Fundamental Law: Memory Bandwidth vs. Compute

During autoregressive token generation (the decoding phase), an LLM generates one token at a time. To generate a single token, the model must read every single weight parameter from memory into the processor cores once.

Token Generation Throughput Formula
Maximum Tokens/Second ≈ Memory Bandwidth (GB/s) ÷ Model Memory Footprint (GB)

For example, an 8B model quantized to Q4 (~5 GB) on an M-series Mac with 200 GB/s unified bandwidth has a theoretical ceiling of 200 ÷ 5 ≈ 40 tokens/second. On standard dual-channel DDR5 system RAM (~80 GB/s), that same model tops out around 16 tokens/second.

2. Comparing Compute & Memory Architectures

Apple Silicon Unified Memory (Metal)

Apple’s M-series architecture (M1 through M4) shares a single high-bandwidth memory pool between the CPU, GPU, and Neural Engine. Ollama uses the Metal API to execute matrix multiplications directly on GPU cores without copying weights across a PCIe bus.

Key Strengths
  • High memory bandwidth (100 to 800+ GB/s).
  • Massive VRAM ceiling (up to 128 GB on Max/Ultra chips).
  • Low idle power consumption (~5–15W).
Key Constraints
  • macOS caps GPU allocation (default ~75% of RAM).
  • Memory is non-upgradeable post-purchase.
  • M1/M2/M3 base chips have fewer GPU cores.
NVIDIA Discrete GPUs (CUDA)

Dedicated NVIDIA GeForce RTX and workstation GPUs feature GDDR6 and GDDR6X memory with industry-leading bandwidth (>900 GB/s on RTX 4090). Tensor Cores provide unmatched raw FLOPs for prompt processing (prefill).

Key Strengths
  • Highest tokens/second generation speed.
  • Near-instant prompt prefill via Tensor Cores.
  • Full FP16 / BF16 hardware tensor support.
The VRAM Cliff
  • Strict hardware VRAM ceiling (8 GB, 12 GB, 16 GB, 24 GB).
  • Offloading layers to CPU RAM across PCIe causes severe latency.
  • Higher thermal dissipation and electrical power draw.
AMD ROCm & Pure CPU Execution

When dedicated VRAM or Metal acceleration is unavailable, Ollama falls back to vectorized CPU instructions (AVX2, AVX-512, AMX). While completely stable, throughput is limited by system RAM bandwidth.

CPU Optimization Rules:
• Set thread count equal to physical cores, not logical hyperthreads.
• Dual-channel or quad-channel RAM configurations significantly improve tokens/sec.
• Keep models under 4B–7B parameters to maintain acceptable interactive speeds (~4–10 tok/s).

3. The Hidden Memory Multiplier: KV Cache

Many developers calculate memory needs purely based on the downloaded model file size. This causes catastrophic OOM crashes during long conversations because the Key-Value (KV) Cache grows linearly with every token in the conversation history.

Context LengthKV Cache Overhead (8B Model)KV Cache Overhead (14B Model)Practical Impact
2,048 tokens~0.25 GB~0.45 GBNegligible; fits comfortably
8,192 tokens~1.0 GB~1.8 GBNoticeable allocation chunk
32,768 tokens~4.0 GB~7.2 GBMay exceed 8GB/16GB VRAM bounds
131,072 tokens~16.0 GB~28.8 GBRequires 32GB+ dedicated unified memory

* Calculations based on standard FP16 KV cache representations across 32 transformer layers. FlashAttention and Q8_0 KV cache quantization reduce these numbers by approximately 50%.

Hack Day Sizing Rules of Thumb

1. Keep a 1.5 GB safety buffer: Always leave at least 1.5 GB of free disk space and system memory unallocated for OS background services.
2. Don’t push beyond 75% RAM on macOS: Metal will refuse to allocate the remaining 25% to GPU operations without manual kernel sysctl flags.
3. Avoid layer splitting across PCIe: If a model doesn’t fit 100% inside your NVIDIA VRAM, pick the next smaller quantization or parameter tier. Partial offloading kills throughput.