Run AI agents on a local LLM with Ollama
Griffin talks to any server that speaks the OpenAI chat completions API — Ollama, LM Studio, vLLM, llama.cpp's server and others. This guide gets a fully local agent running and explains which models work well with tools.
1. Run a model that supports tool calling
Agents call tools, so the model must support function calling. With Ollama, models such as qwen2.5, llama3.1 and mistral-nemo do; bigger variants follow multi-step instructions noticeably better.
ollama pull qwen2.5:14b
ollama serve # listens on http://127.0.0.1:11434
2. Point Griffin at it
Griffin uses the OpenAI-compatible engine when GRIFFIN_OPENAI_BASE_URL is set. A local server usually needs no key.
From source, on the same machine as Ollama:
GRIFFIN_OPENAI_BASE_URL=http://127.0.0.1:11434/v1 \
GRIFFIN_OPENAI_MODEL=qwen2.5:14b \
GRIFFIN_DATA=./data GRIFFIN_WORKSPACE=./workspace PORT=3100 \
node apps/server/src/index.mjs
With Docker, reach Ollama on the host through host.docker.internal:
docker run -d --name griffin -p 3100:3100 -v griffin-data:/data \
--add-host=host.docker.internal:host-gateway \
-e GRIFFIN_OPENAI_BASE_URL=http://host.docker.internal:11434/v1 \
-e GRIFFIN_OPENAI_MODEL=qwen2.5:14b \
ghcr.io/paziresh24/griffin:latest
On a fresh install the built-in agent picks the engine whose settings are present, so with only these two variables it runs on your local model. Other servers work the same way: LM Studio's default is http://127.0.0.1:1234/v1, vLLM's http://127.0.0.1:8000/v1.
3. Mix local and hosted models
Each agent chooses its own engine and model under Agents → engine and model. A practical split is a local model for high-volume, low-risk agents (summaries, note keeping, classification) and a hosted frontier model for the orchestrator that plans and delegates. You can also pick a model for a single chat from the model menu.
Tips for local models
- Fewer tools per agent help small models a lot: give each agent only the tools its job needs.
- Keep instructions short and concrete; small models follow examples better than long rules.
- If tool calls come back as plain text, the model or its chat template doesn't support tools — try another model.
- Watch context length: long tool outputs fill a small context window quickly.
Griffin is open source (MIT): github.com/paziresh24/Griffin · self-hosted agent platform overview.