# Local models: AI that stays on your computer.
> Run Ask AI, the classifiers and document reading on Ollama, LM Studio or any OpenAI-compatible server. Mix local and cloud by role.
[Get started](https://app.getoatmilk.com/sign-up) · [Sign in](https://app.getoatmilk.com/sign-in) · [Docs](https://getoatmilk.com/docs.md) · [llms.txt](https://getoatmilk.com/llms.txt)
## How it works
1. Start Ollama or LM Studio
2. Oatmilk picks models that fit
3. Test each one, then use it
## What it does
- **Ollama, LM Studio or any OpenAI-compatible server.** Ollama, LM Studio, vLLM, llama.cpp and LocalAI all work, with no code changes.
- **Models that fit your memory.** Qwen3.5 9B for chat and reading, or 4B on a computer with less than 14 GB of memory.
- **The classifier, on your machine.** tev1, a decision model served by Ollama 0.35 and later, answers the same questions as TypeSafe's Jev.
- **Nothing leaves this computer.** With every role local, no model call goes out.
- **Local and cloud by role.** Keep classifiers and document reading here while Ask AI uses a cloud model, or the other way round.
- **oatmilk models.** List, download and test the models Oatmilk uses from the terminal.
## Servers
- **Ollama** (`http://localhost:11434/v1`). Use 0.35 or later for decision models (`tev1`, the larger `nimble`), which answer the same typed questions as TypeSafe's Jev.
- **LM Studio** (`http://localhost:1234/v1`). Start its server in the Developer tab or with `lms server start`.
- **Any OpenAI-compatible server**: vLLM, llama.cpp `llama-server`, LocalAI. Set its `/v1` address.
`oatmilk setup` finds Ollama or LM Studio on this computer, picks models that fit, downloads missing ones into Ollama and writes every setting.
## Recommended models
| Role | Ollama | LM Studio |
| --- | --- | --- |
| Chat, tools and reading documents | `qwen3.5:9b` (`qwen3.5:4b` under 14 GB of memory) | `qwen/qwen3.5-9b` (`qwen/qwen3.5-4b` under 14 GB) |
| Classifier | `tev1`, a decision model through Ollama's systemone endpoint | `qwen/qwen3.5-4b`, a chat model |
| Embeddings | `nomic-embed-text` | nomic-embed-text v1.5 |
Pick models that support tool calling and structured output. Setup reads the computer's memory: under 14 GB it picks the 4B chat model. Long receipts and statements need a long context, for example `OLLAMA_CONTEXT_LENGTH=32768 ollama serve`.
## Classifiers and confidence
Oatmilk's classifiers (categories, receipt matching, mail routing, project tags, search ranking) ask typed questions and get a probability for every answer. Locally there are two ways to answer them:
- **A decision model** (`tev1` or `nimble` on Ollama 0.35+, `OATMILK_LOCAL_CLASSIFIER_API=systemone`). Answers carry a confidence in the same place as Jev's, so every step works as it does with Jev and keeps the same gates. Automatic receipt matching still needs 98% probability and 95% confidence.
- **A chat model** (`OATMILK_LOCAL_CLASSIFIER_API=chat`, the default; LM Studio and other servers). It returns the model's own probabilities, which aren't calibrated, and no confidence. Steps that need confidence, such as matching a receipt to a bank line, always wait for a person.
## Mix local and cloud by role
`OATMILK_LOCAL_ROLES` names the calls that run locally: `chat`, `vision`, `classifier`, `embedding`, `transcription`, comma-separated. The others use cloud models through Vercel AI Gateway. Leave it out to run every role locally.
```sh
OATMILK_MODEL_PROVIDER=ollama
OATMILK_LOCAL_MODEL=qwen3.5:9b
OATMILK_LOCAL_CLASSIFIER_MODEL=tev1
OATMILK_LOCAL_CLASSIFIER_API=systemone
OATMILK_LOCAL_ROLES=classifier,vision # classifiers and documents here, Ask AI in the cloud
```
## Settings
| Variable | What it does |
| --- | --- |
| `OATMILK_MODEL_PROVIDER` | `gateway` (the default), `ollama`, `lmstudio` or `openai-compatible`. |
| `OATMILK_LOCAL_BASE_URL` | The server's `/v1` address. Required for `openai-compatible`. |
| `OATMILK_LOCAL_MODEL` | Answers every chat, tool-calling, extraction and review call. |
| `OATMILK_LOCAL_VISION_MODEL` | For prompts with an image or PDF page. Defaults to `OATMILK_LOCAL_MODEL`. |
| `OATMILK_LOCAL_CLASSIFIER_MODEL` | Stands in for Jev in every classifier. |
| `OATMILK_LOCAL_CLASSIFIER_API` | `systemone` for a decision model, `chat` to ask a chat model. |
| `OATMILK_LOCAL_EMBEDDING_MODEL` | For embeddings, for example `nomic-embed-text`. |
| `OATMILK_LOCAL_ROLES` | Which calls run locally; the rest use cloud models. |
| `OATMILK_LOCAL_MODEL_MAP` | Pins particular cloud model ids to particular local models, for example `typesafe-ai/jev=qwen3:4b`. |
| `OATMILK_LOCAL_REASONING` | Caps thinking for every local chat call; `none` answers at once on a laptop. |
An invalid setting stops the server at start-up with a message naming the variable to fix.
## Commands
```sh
oatmilk models # the servers on this computer and the model for each role
oatmilk models pull qwen3.5:9b # download a model into Ollama
oatmilk models test # ask each model to answer
oatmilk models use --models lmstudio --restart # switch servers
```
## Running Oatmilk itself on your computer
Local models work with Oatmilk running on your computer or your own server. Oatmilk runs in Docker and needs about 4 GB of memory and 15 GB of disk, plus what the models need; the first build takes about 10 minutes. See https://getoatmilk.com/local and https://getoatmilk.com/docs/self-hosting.
## More from Oatmilk
- [Classifier](https://getoatmilk.com/classifier.md): Each bank and card line runs through rules, memory, receipts and Jev. Oatmilk sorts it only when it's sure, and asks you one question when it isn't.
- [Evidence matching](https://getoatmilk.com/evidence.md): Oatmilk reads receipts from your company's inboxes and matches each one to its charge. It asks you when it isn't sure and keeps the original.
- [Bookkeeping](https://getoatmilk.com/bookkeeping.md): Bank, card, Wise and Stripe charges come in and get sorted. Oatmilk asks you only when it isn't sure.
- [Compliance](https://getoatmilk.com/stay-compliant.md): Ask any rules question and get an answer with official sources. Guides, deadlines and rule changes for your company, in one place.
- [Invoices](https://getoatmilk.com/invoices.md): Make an invoice in a minute, send it, and see when it's opened and paid.
- [Agreements](https://getoatmilk.com/agreements.md): Write an agreement, send it to sign on any device, and keep every signed copy with a full record.
- [Contractors](https://getoatmilk.com/contractors.md): Contractors send hours from their own portal. You approve with a tap, pay with Wise and get tax slips at year-end.
- [Tax](https://getoatmilk.com/tax.md): Year-end for Canadian companies, one plain question at a time. T2, GST/HST and T4A work, ready for your accountant.
- [Ask AI](https://getoatmilk.com/ask-ai.md): Ask a question about your company and get a straight answer. Ask AI reads your books and asks before it changes anything.
- [Accountants](https://getoatmilk.com/accountants.md): Invite your accountant to see the books and help with year-end. You pick what they can do and when it ends, and every action is on the record.
- [Run it locally](https://getoatmilk.com/local.md): Run Oatmilk on your own computer with the CLI and the web portal, or choose where each part runs: this computer, the cloud, or off.
- [CLI](https://getoatmilk.com/cli.md): Oatmilk in your terminal: every page one keypress away, Ask AI built in, and themes that work in light and dark terminals.
- [AI connectors](https://getoatmilk.com/mcp.md): Connect Claude, ChatGPT, Cursor, Codex and more to Oatmilk. Your AI can do what you can do, and nothing more.
Human version: https://getoatmilk.com/local-models