Hardware-Aware Local AI Recommender
Match your machine’s RAM, GPU/VRAM, and OS against verified open-weight models. Deterministic memory sizing and instant TypeScript starter generation for Ollama.
The Physics of Local LLM Sizing: Why VRAM & Architecture Matter
A common failure mode at AI hackathons is attempting to run a local model that exceeds physical hardware capacity, leading to disk swapping, system lockups, or out-of-memory crashes. Running open-weight models locally requires understanding three distinct memory allocations.
The downloaded GGUF file size represents only the compressed weights. Once loaded into active memory, the KV cache (which grows linearly with context length) and active runtime buffers demand substantial additional allocation.
Unified memory allows GPU cores direct access to system RAM at up to 800 GB/s bandwidth without PCIe transfer penalties. However, macOS caps GPU allocation (typically 75% of total RAM via sysctl iogpu.wired_mem_limit).
Dedicated GDDR6/HBM memory offers ultra-fast bandwidth (>900 GB/s on RTX 4090), but capacity is rigid. When a model exceeds VRAM, offloading layers to CPU RAM across the PCIe bus drops token generation from 40+ tok/s to <5 tok/s.
Without dedicated CUDA or unified Metal memory, inference runs across CPU cores via AVX-512/AMX instructions. Memory bandwidth (typically 50–90 GB/s) bottlenecks throughput to ~3–8 tok/s, making compact 3B–4B models optimal.
Hardware Tier Sizing Matrix
Baseline recommendations based on physical memory tiers and verified Ollama model weights.
| Hardware Tier | Typical Spec | Max Weight Size | Recommended Models | Expected Speed |
|---|---|---|---|---|
| 8 GB Entry | M1/M2/M3 Air, Intel/AMD 8GB RAM | < 3.5 GB | Qwen 3.5 4B, Gemma 4 E4B, DeepSeek R1 1.5B | 12–25 tok/s |
| 16 GB Balanced | MacBook 16GB, RTX 3060/4060 (8-12GB VRAM) | < 6.5 GB | Qwen 3.5 9B, Qwen 3 8B, Phi-4 Mini | 25–45 tok/s |
| 24–36 GB Power | M-Series Pro/Max, RTX 4080 (16GB VRAM) | < 14 GB | Gemma 4 12B, GPT-OSS 20B (MoE) | 35–60 tok/s |
| 64 GB+ Workstation | M-Series Max/Ultra, RTX 3090/4090 (24GB VRAM) | < 28 GB | Qwen 3.6 27B/35B, Qwen 3.8 27B | 40–80 tok/s |
Source-Observed Facts vs. Empirical Estimates
Hack Day Starter strictly separates primary source data from empirical sizing calculations.
Extracted directly from official Ollama library manifests: exact byte sizes, sha256 layer digests, context window token caps, and input modalities (vision, audio).
Officially published system memory guidance directly from foundation model creators (Google DeepMind, Alibaba Cloud, Mistral AI) in model cards.
Empirical memory comfort calculations including KV cache overhead and a 1.5 GB safety buffer, clearly labeled as estimates to prevent developer surprises.
Technical Architecture & Guides
In-depth engineering documentation for running local AI models and building agents.
Browse all 17 tracked entries with exact manifest byte sizes, context windows, lifecycle states, and verified capabilities.
Learn how Apple Silicon Metal unified memory, discrete NVIDIA CUDA VRAM, and KV cache calculations impact performance.
Build autonomous TypeScript agent loops with native Ollama function calling, AST-safe calculators, and zero heavy frameworks.
Explore our mathematical recommendation formula, hard safety gates, 1.5 GB buffer policy, and zero-hallucination design.