Hack Day Starter
Catalog verified:October 4, 2026 13:00 UTCOfficial Ollama Library ↗

Hardware-Aware Local AI Recommender

Match your machine’s RAM, GPU/VRAM, and OS against verified open-weight models. Deterministic memory sizing and instant TypeScript starter generation for Ollama.

Your computer

Configure your operating system and hardware to evaluate model compatibility.

1. Operating system
2. Hardware typeAvailable on macOS
Apple silicon generationM1 through M6 architecture

Note: All Apple silicon generations share the unified memory architecture. Model fit is determined by available unified memory.

Unified memorySelected: 24 GB

Unified memory is shared between CPU, GPU, and neural engine.

Free disk spaceSelected: 50 GB

Evaluated against the exact model download size plus a recommended 1.5 GB safety buffer.

Must support:
Hardware Engineering

The Physics of Local LLM Sizing: Why VRAM & Architecture Matter

A common failure mode at AI hackathons is attempting to run a local model that exceeds physical hardware capacity, leading to disk swapping, system lockups, or out-of-memory crashes. Running open-weight models locally requires understanding three distinct memory allocations.

Total Working Memory Formula
Total Memory Required = Model Weights + KV Cache + Run-Time Context + OS Overhead

The downloaded GGUF file size represents only the compressed weights. Once loaded into active memory, the KV cache (which grows linearly with context length) and active runtime buffers demand substantial additional allocation.

Apple Silicon (Metal)

Unified memory allows GPU cores direct access to system RAM at up to 800 GB/s bandwidth without PCIe transfer penalties. However, macOS caps GPU allocation (typically 75% of total RAM via sysctl iogpu.wired_mem_limit).

NVIDIA Discrete (CUDA)

Dedicated GDDR6/HBM memory offers ultra-fast bandwidth (>900 GB/s on RTX 4090), but capacity is rigid. When a model exceeds VRAM, offloading layers to CPU RAM across the PCIe bus drops token generation from 40+ tok/s to <5 tok/s.

AMD & CPU Execution

Without dedicated CUDA or unified Metal memory, inference runs across CPU cores via AVX-512/AMX instructions. Memory bandwidth (typically 50–90 GB/s) bottlenecks throughput to ~3–8 tok/s, making compact 3B–4B models optimal.

Quick Reference

Hardware Tier Sizing Matrix

Baseline recommendations based on physical memory tiers and verified Ollama model weights.

Hardware TierTypical SpecMax Weight SizeRecommended ModelsExpected Speed
8 GB EntryM1/M2/M3 Air, Intel/AMD 8GB RAM< 3.5 GBQwen 3.5 4B, Gemma 4 E4B, DeepSeek R1 1.5B12–25 tok/s
16 GB BalancedMacBook 16GB, RTX 3060/4060 (8-12GB VRAM)< 6.5 GBQwen 3.5 9B, Qwen 3 8B, Phi-4 Mini25–45 tok/s
24–36 GB PowerM-Series Pro/Max, RTX 4080 (16GB VRAM)< 14 GBGemma 4 12B, GPT-OSS 20B (MoE)35–60 tok/s
64 GB+ WorkstationM-Series Max/Ultra, RTX 3090/4090 (24GB VRAM)< 28 GBQwen 3.6 27B/35B, Qwen 3.8 27B40–80 tok/s
Data Integrity

Source-Observed Facts vs. Empirical Estimates

Hack Day Starter strictly separates primary source data from empirical sizing calculations.

Source-Observed Data

Extracted directly from official Ollama library manifests: exact byte sizes, sha256 layer digests, context window token caps, and input modalities (vision, audio).

Vendor Guidance

Officially published system memory guidance directly from foundation model creators (Google DeepMind, Alibaba Cloud, Mistral AI) in model cards.

Starter Estimates

Empirical memory comfort calculations including KV cache overhead and a 1.5 GB safety buffer, clearly labeled as estimates to prevent developer surprises.

Explore Guides

Technical Architecture & Guides

In-depth engineering documentation for running local AI models and building agents.

Hack Day Starter • Built for Hacktoberfest 2026 — Weekend Challenge: Build for a Friend

Verified Model Registry v2026.10.04.1 (October 4, 2026 13:00 UTC)