How to connect local Ollama models
Connect a local Ollama server to your BYOK chat workspace
Ollama runs open-weight models on your own machine and exposes an OpenAI-compatible endpoint. Connect it to MultimodelChat with a local base URL — no API key or per-token cost, since the model runs on your hardware.
Not requiredLocal models have no per-token cost — you pay only for your own hardware and electricity.
How to connect a local Ollama model
Install Ollama, pull a model, and make sure the server is running. Ollama exposes an OpenAI-compatible endpoint at http://localhost:11434/v1, which you connect in MultimodelChat as an OpenAI-Compatible provider.
- Install Ollama for your operating system
- Pull at least one model
- Confirm the Ollama server is running
- Connect base URL http://localhost:11434/v1 (use a tunnel if MultimodelChat is cloud-hosted)
Connect it to MultimodelChat
Once you have the required credentials, start your free trial, add this provider in Settings, and save the connection. MultimodelChat encrypts the API key before storing it.
Available models on Multimodel Chat
Any model you have pulled in Ollama is available. Use the model name you pulled (for example llama3.2) as the model ID.
Step-by-step guide
Install Ollama
Download and install for your platform.
- Go to https://ollama.com/download
- Install the app for macOS, Windows, or Linux
- Confirm the command line tool is available by running ollama --version
Pull a model
Download the weights you want to chat with.
- Run a pull command, for example: ollama pull llama3.2
- Wait for the download to finish
- Note the exact model name — it is the model ID you will use
Make sure the server is running
Ollama serves requests on a local port.
- The desktop app starts the server automatically
- On a headless machine, run: ollama serve
- The endpoint listens on http://localhost:11434 by default
Connect it as an OpenAI-Compatible provider
Ollama exposes an OpenAI-compatible endpoint.
- In MultimodelChat, go to Settings → API Keys
- Select "OpenAI-Compatible"
- Base URL: http://localhost:11434/v1
- API key: any value (Ollama ignores it locally)
- Model ID: the model you pulled, e.g. llama3.2
- Save — MultimodelChat tests the local connection
If MultimodelChat is cloud-hosted, localhost points at the server, not your machine. Expose Ollama through a tunnel or connect over your network so the app can reach it. See our local-models guide for the full walkthrough.
Common issues
Swipe horizontally to see the full table.
| Issue | Fix |
|---|---|
| Connection refused | Start the Ollama server (ollama serve) and confirm it listens on port 11434 |
| Model not found | Pull the model first (ollama pull <model>) and use its exact name |
| Cloud app cannot reach localhost | Use a tunnel (for example ngrok) or a reachable host, then use that URL as the base URL |
| CORS errors | Set OLLAMA_ORIGINS to allow your app's origin when calling from a browser |
Frequently asked questions
Does a local model cost anything?
No per-token charge — Ollama runs the model on your own machine, so the only cost is your hardware and power.
Can I use Ollama with the cloud app?
Yes, if the app can reach the server. Because localhost on a cloud-hosted app refers to the server, expose Ollama through a tunnel or your network and use that URL as the base URL.
Set up another provider
Related guides
Provider requirements and pricing last reviewed October 10, 2026. Check the linked provider documentation for current terms before purchasing credits.
Ready to connect your Ollama (local) key?
Paste your API key and Multimodel Chat tests it live against the provider before saving. The key is encrypted in this browser's vault and used only for your requests.
Start the 7-day free trialNew to the workflow? See how it works in three steps