Building Local Tool-Calling Agents with Ollama
How to implement native function calling and multi-step autonomous agent loops in pure TypeScript without heavyweight agent frameworks.
1. How Native Tool Calling Works in Ollama
Unlike early prompt-engineering approaches where models were asked to format outputs as markdown code blocks, modern open-weight models (Qwen 3.5, Gemma 4, Llama 3.1) support native function calling directly via Ollama’s /api/chat endpoint.
When you pass a tools array containing standard JSON Schema definitions, Ollama utilizes grammar-constrained decoding to force the model to output syntactically valid function parameters.
2. The Multi-Step Agent Loop Architecture
3. Security: The Dangers of eval() in Local Agents
Many naive tutorials evaluate math expressions or code emitted by LLMs using JavaScript’s eval() or new Function(). This creates critical Remote Code Execution (RCE) vulnerabilities if prompt injection or untrusted inputs manipulate the model into executing arbitrary system commands.
// DANGEROUS: Arbitrary code execution vulnerability const answer = eval(toolCall.arguments.expression);
Hack Day Starter’s generated agent templates implement a pure recursive-descent math tokenizer that only parses numbers, parentheses, and arithmetic operators (+, -, *, /, ^, %). Zero eval, zero Function constructors, 100% auditable.
Verified Tool-Calling Models (19)
Current open-weight models with verified function callingGoogle's 4B multimodal model with native tool-calling, vision inputs, and configurable reasoning tokens.
ollama pull gemma4:e4bFlagship local model from Google offering high reasoning, multimodal perception, and 256k context.
ollama pull gemma4:12bCompact Qwen 3.5 release featuring hybrid thinking tokens, strong code logic, vision input, and 256k context.
ollama pull qwen3.5:4bBalanced 9B model excelling at complex multi-step reasoning, coding tasks, and 256k context window.
ollama pull qwen3.5:9bDense 27B model from Qwen 3.6 series with extended 256k context window and multimodal vision.
ollama pull qwen3.6:27b35B parameter model with 23 GB download footprint and full 256k context for high-memory setups.
ollama pull qwen3.6:35bQwen 3.8 release (August 2026) with refined reasoning architecture, 256k context, and multimodal vision.
ollama pull qwen3.8:27bStandard dense 8B model with fast latency and proven tool calling for intermediate laptops.
ollama pull qwen3:8bOpen-weight architecture with 21B total parameters (3.6B active) and 14 GB Ollama download.
ollama pull gpt-oss:20bMicrosoft's 3.8B model with strong synthetic math, reasoning, and tool support for lightweight machines.
ollama pull phi4-miniNVIDIA's 4B edge model with 2.8 GB footprint and 256k context engineered for low-latency tool calling.
ollama pull nemotron-3-nano:4b80B sparse MoE coding model with 3B active parameters, 52 GB footprint, and 256k context for high-end workstations.
ollama pull qwen3-coder-next:latestCandidate model entry pending full verification audit.
ollama pull glm4:9bMeta's 3B lightweight model. Runtime verified live through Hack Day Starter Chat + Tool-calling Agent.
ollama pull llama3.2:3bMeta's foundational 8B model with 128k context. Runtime verified live through Hack Day Starter Chat + Tool-calling Agent.
ollama pull llama3.1:8bAlibaba's dedicated 7B code generation model with 32k context. Source/capability verified; runtime verification pending.
ollama pull qwen2.5-coder:7b12B multilingual model with Tekken tokenizer and 1000k context. Source/capability verified; runtime verification pending.
ollama pull mistral-nemo:12bLiquid Foundation Model with reasoning tokens and 125k context. Source/capability verified; runtime verification pending.
ollama pull lfm2.5:8b24B multimodal code agent model with 384k context. Source/capability verified; runtime verification pending.
ollama pull devstral-small-2:latest