You use three different model providers, each with its own chat window, subscription and history. Open WebUI puts them behind one interface you host yourself: pick a model per conversation, keep all history on your server, share it with your team, and plug in local models when you want nothing to leave the machine.
Two ways to run it
As a front end for APIs. Open WebUI connects to any OpenAI-compatible API — hosted providers, a LiteLLM proxy, or your own endpoint. The server does almost no work; a small plan is plenty.
With local models. Add Ollama on the same server and run open-weight models locally. Now RAM is the constraint, and CPU inference is slow. Honest trade-off: total privacy, modest speed.
| Setup | Plan |
|---|---|
| API front end, personal use | Micro — 1 vCPU, 2 GB, $5/mo (SSH tunnel) |
| API front end, team access over HTTPS | Micro-IP — 1 vCPU, 2 GB, $10/mo |
| Local 3B models via Ollama | AI-Agent — 2 vCPU, 4 GB, $10/mo |
| Local 7–8B models | Medium — 4 vCPU, 6 GB, $12/mo, or Pro plans for bigger ones |
Install with Docker
On a server with Docker installed:
docker run -d --name open-webui --restart always \
-p 127.0.0.1:3000:8080 \
-v open-webui:/app/backend/data \
ghcr.io/open-webui/open-webui:main
Binding to 127.0.0.1 keeps it off the public internet until you decide how to expose it. Chats, users and settings live in the open-webui volume.
Access it safely
SSH tunnel — works on every plan, including NAT:
ssh -p 22 -N -L 3000:127.0.0.1:3000 you@203.0.113.10
Open http://localhost:3000. On a NAT plan, use your personal SSH port.
HTTPS on your domain — for a team on a dedicated-IP plan. Point a subdomain at the server (guide) and proxy it with nginx, including WebSocket headers:
location / {
proxy_pass http://127.0.0.1:3000;
proxy_set_header Host $host;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection upgrade;
proxy_read_timeout 300s;
}
Then run certbot --nginx -d chat.example.com. The long read timeout matters: slow models stream answers for minutes.
First steps inside
- Register immediately. The first account becomes the admin. Don't leave a fresh install public — someone else could claim it.
- Close sign-ups. Admin settings → set the default role for new users to pending, or disable registration.
- Add a connection. Settings → Connections → paste an API base URL and key. Models from that provider appear in the model picker.
Adding local models
Install Ollama on the host (our guide), pull a small model, and start Open WebUI so it can reach it:
ollama pull llama3.2:3b
docker rm -f open-webui
docker run -d --name open-webui --restart always \
-p 127.0.0.1:3000:8080 --add-host=host.docker.internal:host-gateway \
-e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
-v open-webui:/app/backend/data ghcr.io/open-webui/open-webui:main
The volume keeps your chats across the restart. Expect a few tokens per second on CPU — fine for summaries and private notes, slow for long essays. For heavier local inference, see VPS for Ollama.
Worth knowing
- Updates are frequent.
docker pull ghcr.io/open-webui/open-webui:main, then recreate the container. Back up the volume first. - API costs are yours. Open WebUI doesn't limit usage per user by default; watch spending if you share it with a team.
- It's a chat app, not a model. Answer quality depends entirely on the model behind it.
Plan limits and upgrades are in the plans docs.
Comments
No comments yet. Be the first.