Hack Day Starter
Agent Architecture GuideTypeScript & Ollama

Building Local Tool-Calling Agents with Ollama

How to implement native function calling and multi-step autonomous agent loops in pure TypeScript without heavyweight agent frameworks.

1. How Native Tool Calling Works in Ollama

Unlike early prompt-engineering approaches where models were asked to format outputs as markdown code blocks, modern open-weight models (Qwen 3.5, Gemma 4, Llama 3.1) support native function calling directly via Ollama’s /api/chat endpoint.

When you pass a tools array containing standard JSON Schema definitions, Ollama utilizes grammar-constrained decoding to force the model to output syntactically valid function parameters.

2. The Multi-Step Agent Loop Architecture

// 1. User submits task
messages = [{ role: "user", content: "What is 482 * 19.5?" }]
↓
// 2. Ollama evaluates available tools
POST /api/chat with { model, messages, tools: [calculatorTool] }
↓
// 3. Model requests tool call
response.tool_calls = [{ function: { name: "calculate", arguments: { expr: "482 * 19.5" } } }]
↓
// 4. Client executes function locally
result = executeCalculator("482 * 19.5") // 9399
↓
// 5. Result appended with role: "tool"
messages.push({ role: "tool", content: "9399" })
↓
// 6. Model synthesizes final human response
"482 multiplied by 19.5 equals 9,399."

3. Security: The Dangers of eval() in Local Agents

Many naive tutorials evaluate math expressions or code emitted by LLMs using JavaScript’s eval() or new Function(). This creates critical Remote Code Execution (RCE) vulnerabilities if prompt injection or untrusted inputs manipulate the model into executing arbitrary system commands.

Insecure Pattern (Never Do This):
// DANGEROUS: Arbitrary code execution vulnerability
const answer = eval(toolCall.arguments.expression);
Hack Day Starter Approach (AST / Recursive-Descent):

Hack Day Starter’s generated agent templates implement a pure recursive-descent math tokenizer that only parses numbers, parentheses, and arithmetic operators (+, -, *, /, ^, %). Zero eval, zero Function constructors, 100% auditable.

Verified Tool-Calling Models (19)

Current open-weight models with verified function calling
Gemma 4 e4b4B

Google's 4B multimodal model with native tool-calling, vision inputs, and configurable reasoning tokens.

ollama pull gemma4:e4b
Gemma 4 12B12B

Flagship local model from Google offering high reasoning, multimodal perception, and 256k context.

ollama pull gemma4:12b
Qwen 3.5 4B4B

Compact Qwen 3.5 release featuring hybrid thinking tokens, strong code logic, vision input, and 256k context.

ollama pull qwen3.5:4b
Qwen 3.5 9B9B

Balanced 9B model excelling at complex multi-step reasoning, coding tasks, and 256k context window.

ollama pull qwen3.5:9b
Qwen 3.6 27B27B

Dense 27B model from Qwen 3.6 series with extended 256k context window and multimodal vision.

ollama pull qwen3.6:27b
Qwen 3.6 35B35B

35B parameter model with 23 GB download footprint and full 256k context for high-memory setups.

ollama pull qwen3.6:35b
Qwen 3.8 27B27B

Qwen 3.8 release (August 2026) with refined reasoning architecture, 256k context, and multimodal vision.

ollama pull qwen3.8:27b
Qwen 3 8B8B

Standard dense 8B model with fast latency and proven tool calling for intermediate laptops.

ollama pull qwen3:8b
GPT-OSS 20B21B

Open-weight architecture with 21B total parameters (3.6B active) and 14 GB Ollama download.

ollama pull gpt-oss:20b
Phi-4 Mini3.8B

Microsoft's 3.8B model with strong synthetic math, reasoning, and tool support for lightweight machines.

ollama pull phi4-mini
Nemotron 3 Nano 4B4B

NVIDIA's 4B edge model with 2.8 GB footprint and 256k context engineered for low-latency tool calling.

ollama pull nemotron-3-nano:4b
Qwen 3 Coder Next80B

80B sparse MoE coding model with 3B active parameters, 52 GB footprint, and 256k context for high-end workstations.

ollama pull qwen3-coder-next:latest
GLM-4 9B (Unverified Candidate)9B

Candidate model entry pending full verification audit.

ollama pull glm4:9b
Llama 3.2 3B3B

Meta's 3B lightweight model. Runtime verified live through Hack Day Starter Chat + Tool-calling Agent.

ollama pull llama3.2:3b
Llama 3.1 8B8B

Meta's foundational 8B model with 128k context. Runtime verified live through Hack Day Starter Chat + Tool-calling Agent.

ollama pull llama3.1:8b
Qwen 2.5 Coder 7B7B

Alibaba's dedicated 7B code generation model with 32k context. Source/capability verified; runtime verification pending.

ollama pull qwen2.5-coder:7b
Mistral NeMo 12B12B

12B multilingual model with Tekken tokenizer and 1000k context. Source/capability verified; runtime verification pending.

ollama pull mistral-nemo:12b
LFM 2.5 8B8B

Liquid Foundation Model with reasoning tokens and 125k context. Source/capability verified; runtime verification pending.

ollama pull lfm2.5:8b
Devstral Small 224B

24B multimodal code agent model with 384k context. Source/capability verified; runtime verification pending.

ollama pull devstral-small-2:latest