Apple’s M-series architecture (M1 through M4) shares a single high-bandwidth memory pool between the CPU, GPU, and Neural Engine. Ollama uses the Metal API to execute matrix multiplications directly on GPU cores without copying weights across a PCIe bus.
- High memory bandwidth (100 to 800+ GB/s).
- Massive VRAM ceiling (up to 128 GB on Max/Ultra chips).
- Low idle power consumption (~5–15W).
- macOS caps GPU allocation (default ~75% of RAM).
- Memory is non-upgradeable post-purchase.
- M1/M2/M3 base chips have fewer GPU cores.