Self-hosting
Use local AI models
Run every model Oatmilk uses on your own hardware with Ollama, LM Studio, vLLM or llama.cpp.
Oatmilk uses AI models to read statements, receipts and documents, to sort mail and match receipts, and for Ask AI. By default it calls hosted models through Vercel AI Gateway. Pointed at a model server on your own hardware instead, every one of those calls stays with you, nothing is billed, and no API key is needed.
This works with Ollama, LM Studio and any server that speaks the OpenAI chat completions API, such as vLLM, llama.cpp's llama-server and LocalAI.
What your computer needs
A rough guide to the memory the models need, on top of Oatmilk's own 4 GB:
| Setup | Memory | Good for |
|---|---|---|
| A 9B chat model and a 4B classifier | 16 GB, more is better | Real books on one computer |
| 4B models | 8 GB | Trying it out; slower and less accurate |
| 0.8B models | 4 GB | Checking that everything is connected, not real use |
| 30B models and larger, on a GPU server | 32 GB of GPU memory or more | A team, faster and more accurate |
A Mac with Apple silicon, or an NVIDIA GPU, makes answers much faster than a CPU alone.
Option 1: Ollama
Use Ollama 0.35 or newer, which adds decision models that answer Oatmilk's classifier questions with a confidence, as the hosted classifier does.
1. Install Ollama and pull the models
ollama pull qwen3.5:9b # chat, tools, extraction and receipt images
ollama pull tev1 # classifiers: a decision model (nimble is larger and more accurate)2. Start it with a longer context
Ollama loads models with a short context by default. Long statements and the agents need more:
OLLAMA_CONTEXT_LENGTH=32768 ollama serveOn Linux, Oatmilk's container reaches your computer through Docker's network, not 127.0.0.1, so Ollama must listen on every address:
OLLAMA_HOST=0.0.0.0 OLLAMA_CONTEXT_LENGTH=32768 ollama serveWith Ollama's systemd service, set both with sudo systemctl edit ollama instead. Docker Desktop on macOS and Windows reaches Ollama either way.
3. Point Oatmilk at it
In bun run self-host setup, choose Ollama on this computer, and keep qwen3.5:9b and tev1 as the models. Setup writes:
OATMILK_MODEL_PROVIDER=ollama
OATMILK_LOCAL_BASE_URL=http://host.docker.internal:11434/v1
OATMILK_LOCAL_MODEL=qwen3.5:9b
OATMILK_LOCAL_CLASSIFIER_MODEL=tev1
OATMILK_LOCAL_CLASSIFIER_API=systemoneOn an older Ollama, remove the last line, and the chat model answers the classifier questions too.
Option 2: LM Studio
- Download a model in LM Studio, for example Qwen3.5 9B (it reads images too) and, for classifiers, Qwen3.5 4B. Set each one's context length to 32768 or more when you load it.
- Open the Developer tab and start the server, or run
lms server start. It listens on port 1234. - On Linux, turn on Serve on Local Network in the server settings, so Oatmilk's container can reach it.
- In
bun run self-host setup, choose LM Studio on this computer, with the model identifiers LM Studio shows.
OATMILK_MODEL_PROVIDER=lmstudio
OATMILK_LOCAL_BASE_URL=http://host.docker.internal:1234/v1
OATMILK_LOCAL_MODEL=qwen/qwen3.5-9b
OATMILK_LOCAL_CLASSIFIER_MODEL=qwen/qwen3.5-4b
OATMILK_LOCAL_CLASSIFIER_API=chatOption 3: vLLM, llama.cpp or another server
Any server with an OpenAI-compatible /v1 address works, on this computer or another one on your network. For example, vLLM on a GPU server:
vllm serve Qwen/Qwen3-8B --enable-auto-tool-choice --tool-call-parser hermes --max-model-len 32768In setup, choose Another OpenAI-compatible server and give its address, ending in /v1, and the model's name. An address typed as localhost is pointed back at your computer from inside the container.
OATMILK_MODEL_PROVIDER=openai-compatible
OATMILK_LOCAL_BASE_URL=http://192.168.1.50:8000/v1
OATMILK_LOCAL_MODEL=Qwen/Qwen3-8B
OATMILK_LOCAL_API_KEY= # only if your server asks for oneChoose a model that supports tool calling and structured output, and start the server with tool calling turned on.
Change models later
Edit self-host/.env, then apply it:
bun run self-host up
bun run self-host doctor # checks that the app container reaches the model serverSettings
| Setting | Default | What it does |
|---|---|---|
OATMILK_MODEL_PROVIDER | gateway | gateway, ollama, lmstudio or openai-compatible |
OATMILK_LOCAL_BASE_URL | Ollama's or LM Studio's usual address | The server's /v1 address. Required for openai-compatible. |
OATMILK_LOCAL_API_KEY | none | Sent as a bearer token when set |
OATMILK_LOCAL_MODEL | required | Answers every chat, tool, extraction and review call |
OATMILK_LOCAL_VISION_MODEL | the chat model | Used whenever a prompt has an image, a PDF page or another file |
OATMILK_LOCAL_REASONING | each call's own | Caps thinking: none, minimal, low, medium, high or xhigh. none answers fastest on a laptop. |
OATMILK_LOCAL_CLASSIFIER_MODEL | the chat model | Answers classifier questions: document kinds, mail routing, receipt matches, categories |
OATMILK_LOCAL_CLASSIFIER_API | chat | systemone for an Ollama decision model, chat to ask a chat model |
OATMILK_LOCAL_CLASSIFIER_BASE_URL | OATMILK_LOCAL_BASE_URL | When classifiers run on another server, such as Ollama next to LM Studio |
OATMILK_LOCAL_EMBEDDING_MODEL | none | For embeddings, for example nomic-embed-text |
OATMILK_LOCAL_TRANSCRIPTION_MODEL | none | Voice input in Ask AI, from a server with /audio/transcriptions such as a whisper.cpp server |
OATMILK_LOCAL_TRANSCRIPTION_BASE_URL | OATMILK_LOCAL_BASE_URL | When transcription runs on another server |
OATMILK_LOCAL_MODEL_MAP | none | Pins particular hosted model ids to particular local models, such as typesafe-ai/jev=qwen3:4b |
An invalid setting stops Oatmilk at start-up with a message that names the setting to fix.
Classifiers and automatic matching
Oatmilk's classifiers answer typed questions with probabilities. An Ollama decision model (systemone) also reports a confidence for each answer, the way the hosted classifier does, so automatic steps such as matching a receipt to a bank line keep working: they still need 98% probability and 95% confidence.
A chat model (chat) gives its own estimate of the probabilities but no confidence. Every classifier still works, but steps that need confidence, such as automatic receipt matching, leave the decision to a person.
Which models to choose
| Role | Ollama | LM Studio |
|---|---|---|
| Chat, tools and extraction | qwen3.5:9b (smaller: qwen3.5:4b, gemma4:e4b) | qwen/qwen3.5-9b, google/gemma-4-e4b |
| Classifiers | nimble or tev1, with systemone | qwen/qwen3.5-4b, with chat |
| Receipts and documents (vision) | qwen3.5:9b, gemma4:12b | qwen/qwen3.5-9b, google/gemma-4-12b |
| Embeddings | nomic-embed-text | nomic-embed-text v1.5 |
Qwen3.5 and similar models think before they answer. On slow hardware, set OATMILK_LOCAL_REASONING=none.
Limits
- Quality depends on the model. Check a few real statements and receipts before you trust a model with your books.
- Image generation isn't available locally, and voice input needs a transcription server.
- Some agents look up their model's context window in a public catalog. If you run fully offline and an agent doesn't start, pin its model to a local one with
OATMILK_LOCAL_MODEL_MAPand checkbun run self-host logs app.
Without the Docker stack
The same settings work for bun run dev in .env.local. There, Oatmilk runs on your computer itself, so use http://localhost:11434/v1 (Ollama) or http://localhost:1234/v1 (LM Studio), which are also the defaults. See Develop Oatmilk locally.