Self-hosting

Use local AI models

Run every model Oatmilk uses on your own hardware with Ollama, LM Studio, vLLM or llama.cpp.

15 minutes · Intermediate

Oatmilk uses AI models to read statements, receipts and documents, to sort mail and match receipts, and for Ask AI. By default it calls hosted models through Vercel AI Gateway. Pointed at a model server on your own hardware instead, every one of those calls stays with you, nothing is billed, and no API key is needed.

This works with Ollama, LM Studio and any server that speaks the OpenAI chat completions API, such as vLLM, llama.cpp's llama-server and LocalAI.

What your computer needs

A rough guide to the memory the models need, on top of Oatmilk's own 4 GB:

SetupMemoryGood for
A 9B chat model and a 4B classifier16 GB, more is betterReal books on one computer
4B models8 GBTrying it out; slower and less accurate
0.8B models4 GBChecking that everything is connected, not real use
30B models and larger, on a GPU server32 GB of GPU memory or moreA team, faster and more accurate

A Mac with Apple silicon, or an NVIDIA GPU, makes answers much faster than a CPU alone.

Option 1: Ollama

Use Ollama 0.35 or newer, which adds decision models that answer Oatmilk's classifier questions with a confidence, as the hosted classifier does.

1. Install Ollama and pull the models

Shell
ollama pull qwen3.5:9b   # chat, tools, extraction and receipt images
ollama pull tev1         # classifiers: a decision model (nimble is larger and more accurate)

2. Start it with a longer context

Ollama loads models with a short context by default. Long statements and the agents need more:

Shell
OLLAMA_CONTEXT_LENGTH=32768 ollama serve

On Linux, Oatmilk's container reaches your computer through Docker's network, not 127.0.0.1, so Ollama must listen on every address:

Shell
OLLAMA_HOST=0.0.0.0 OLLAMA_CONTEXT_LENGTH=32768 ollama serve

With Ollama's systemd service, set both with sudo systemctl edit ollama instead. Docker Desktop on macOS and Windows reaches Ollama either way.

3. Point Oatmilk at it

In bun run self-host setup, choose Ollama on this computer, and keep qwen3.5:9b and tev1 as the models. Setup writes:

self-host/.env
OATMILK_MODEL_PROVIDER=ollama
OATMILK_LOCAL_BASE_URL=http://host.docker.internal:11434/v1
OATMILK_LOCAL_MODEL=qwen3.5:9b
OATMILK_LOCAL_CLASSIFIER_MODEL=tev1
OATMILK_LOCAL_CLASSIFIER_API=systemone

On an older Ollama, remove the last line, and the chat model answers the classifier questions too.

Option 2: LM Studio

  1. Download a model in LM Studio, for example Qwen3.5 9B (it reads images too) and, for classifiers, Qwen3.5 4B. Set each one's context length to 32768 or more when you load it.
  2. Open the Developer tab and start the server, or run lms server start. It listens on port 1234.
  3. On Linux, turn on Serve on Local Network in the server settings, so Oatmilk's container can reach it.
  4. In bun run self-host setup, choose LM Studio on this computer, with the model identifiers LM Studio shows.
self-host/.env
OATMILK_MODEL_PROVIDER=lmstudio
OATMILK_LOCAL_BASE_URL=http://host.docker.internal:1234/v1
OATMILK_LOCAL_MODEL=qwen/qwen3.5-9b
OATMILK_LOCAL_CLASSIFIER_MODEL=qwen/qwen3.5-4b
OATMILK_LOCAL_CLASSIFIER_API=chat

Option 3: vLLM, llama.cpp or another server

Any server with an OpenAI-compatible /v1 address works, on this computer or another one on your network. For example, vLLM on a GPU server:

Shell
vllm serve Qwen/Qwen3-8B --enable-auto-tool-choice --tool-call-parser hermes --max-model-len 32768

In setup, choose Another OpenAI-compatible server and give its address, ending in /v1, and the model's name. An address typed as localhost is pointed back at your computer from inside the container.

self-host/.env
OATMILK_MODEL_PROVIDER=openai-compatible
OATMILK_LOCAL_BASE_URL=http://192.168.1.50:8000/v1
OATMILK_LOCAL_MODEL=Qwen/Qwen3-8B
OATMILK_LOCAL_API_KEY=            # only if your server asks for one

Choose a model that supports tool calling and structured output, and start the server with tool calling turned on.

Change models later

Edit self-host/.env, then apply it:

Shell
bun run self-host up
bun run self-host doctor   # checks that the app container reaches the model server

Settings

SettingDefaultWhat it does
OATMILK_MODEL_PROVIDERgatewaygateway, ollama, lmstudio or openai-compatible
OATMILK_LOCAL_BASE_URLOllama's or LM Studio's usual addressThe server's /v1 address. Required for openai-compatible.
OATMILK_LOCAL_API_KEYnoneSent as a bearer token when set
OATMILK_LOCAL_MODELrequiredAnswers every chat, tool, extraction and review call
OATMILK_LOCAL_VISION_MODELthe chat modelUsed whenever a prompt has an image, a PDF page or another file
OATMILK_LOCAL_REASONINGeach call's ownCaps thinking: none, minimal, low, medium, high or xhigh. none answers fastest on a laptop.
OATMILK_LOCAL_CLASSIFIER_MODELthe chat modelAnswers classifier questions: document kinds, mail routing, receipt matches, categories
OATMILK_LOCAL_CLASSIFIER_APIchatsystemone for an Ollama decision model, chat to ask a chat model
OATMILK_LOCAL_CLASSIFIER_BASE_URLOATMILK_LOCAL_BASE_URLWhen classifiers run on another server, such as Ollama next to LM Studio
OATMILK_LOCAL_EMBEDDING_MODELnoneFor embeddings, for example nomic-embed-text
OATMILK_LOCAL_TRANSCRIPTION_MODELnoneVoice input in Ask AI, from a server with /audio/transcriptions such as a whisper.cpp server
OATMILK_LOCAL_TRANSCRIPTION_BASE_URLOATMILK_LOCAL_BASE_URLWhen transcription runs on another server
OATMILK_LOCAL_MODEL_MAPnonePins particular hosted model ids to particular local models, such as typesafe-ai/jev=qwen3:4b

An invalid setting stops Oatmilk at start-up with a message that names the setting to fix.

Classifiers and automatic matching

Oatmilk's classifiers answer typed questions with probabilities. An Ollama decision model (systemone) also reports a confidence for each answer, the way the hosted classifier does, so automatic steps such as matching a receipt to a bank line keep working: they still need 98% probability and 95% confidence.

A chat model (chat) gives its own estimate of the probabilities but no confidence. Every classifier still works, but steps that need confidence, such as automatic receipt matching, leave the decision to a person.

Which models to choose

RoleOllamaLM Studio
Chat, tools and extractionqwen3.5:9b (smaller: qwen3.5:4b, gemma4:e4b)qwen/qwen3.5-9b, google/gemma-4-e4b
Classifiersnimble or tev1, with systemoneqwen/qwen3.5-4b, with chat
Receipts and documents (vision)qwen3.5:9b, gemma4:12bqwen/qwen3.5-9b, google/gemma-4-12b
Embeddingsnomic-embed-textnomic-embed-text v1.5

Qwen3.5 and similar models think before they answer. On slow hardware, set OATMILK_LOCAL_REASONING=none.

Limits

  • Quality depends on the model. Check a few real statements and receipts before you trust a model with your books.
  • Image generation isn't available locally, and voice input needs a transcription server.
  • Some agents look up their model's context window in a public catalog. If you run fully offline and an agent doesn't start, pin its model to a local one with OATMILK_LOCAL_MODEL_MAP and check bun run self-host logs app.

Without the Docker stack

The same settings work for bun run dev in .env.local. There, Oatmilk runs on your computer itself, so use http://localhost:11434/v1 (Ollama) or http://localhost:1234/v1 (LM Studio), which are also the defaults. See Develop Oatmilk locally.