−25%

on annual Windows plans, until 31 Oct. See plans

EQVPS

VPS for Open WebUI: a private chat interface for LLMs

Host Open WebUI on a VPS as a private chat front end for any OpenAI-compatible API or local Ollama models: Docker setup, safe access and sizing.

You use three different model providers, each with its own chat window, subscription and history. Open WebUI puts them behind one interface you host yourself: pick a model per conversation, keep all history on your server, share it with your team, and plug in local models when you want nothing to leave the machine.

Two ways to run it

As a front end for APIs. Open WebUI connects to any OpenAI-compatible API — hosted providers, a LiteLLM proxy, or your own endpoint. The server does almost no work; a small plan is plenty.

With local models. Add Ollama on the same server and run open-weight models locally. Now RAM is the constraint, and CPU inference is slow. Honest trade-off: total privacy, modest speed.

SetupPlan
API front end, personal useMicro — 1 vCPU, 2 GB, $5/mo (SSH tunnel)
API front end, team access over HTTPSMicro-IP — 1 vCPU, 2 GB, $10/mo
Local 3B models via OllamaAI-Agent — 2 vCPU, 4 GB, $10/mo
Local 7–8B modelsMedium — 4 vCPU, 6 GB, $12/mo, or Pro plans for bigger ones

Install with Docker

On a server with Docker installed:

docker run -d --name open-webui --restart always \
  -p 127.0.0.1:3000:8080 \
  -v open-webui:/app/backend/data \
  ghcr.io/open-webui/open-webui:main

Binding to 127.0.0.1 keeps it off the public internet until you decide how to expose it. Chats, users and settings live in the open-webui volume.

Access it safely

SSH tunnel — works on every plan, including NAT:

ssh -p 22 -N -L 3000:127.0.0.1:3000 you@203.0.113.10

Open http://localhost:3000. On a NAT plan, use your personal SSH port.

HTTPS on your domain — for a team on a dedicated-IP plan. Point a subdomain at the server (guide) and proxy it with nginx, including WebSocket headers:

location / {
    proxy_pass http://127.0.0.1:3000;
    proxy_set_header Host $host;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection upgrade;
    proxy_read_timeout 300s;
}

Then run certbot --nginx -d chat.example.com. The long read timeout matters: slow models stream answers for minutes.

First steps inside

  1. Register immediately. The first account becomes the admin. Don't leave a fresh install public — someone else could claim it.
  2. Close sign-ups. Admin settings → set the default role for new users to pending, or disable registration.
  3. Add a connection. Settings → Connections → paste an API base URL and key. Models from that provider appear in the model picker.

Adding local models

Install Ollama on the host (our guide), pull a small model, and start Open WebUI so it can reach it:

ollama pull llama3.2:3b
docker rm -f open-webui
docker run -d --name open-webui --restart always \
  -p 127.0.0.1:3000:8080 --add-host=host.docker.internal:host-gateway \
  -e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
  -v open-webui:/app/backend/data ghcr.io/open-webui/open-webui:main

The volume keeps your chats across the restart. Expect a few tokens per second on CPU — fine for summaries and private notes, slow for long essays. For heavier local inference, see VPS for Ollama.

Worth knowing

Plan limits and upgrades are in the plans docs.

Ready to deploy? Pay with crypto, no KYC — live in about a minute.

Deploy now →

FAQ

How much RAM does Open WebUI need?

Open WebUI itself runs in about 1 GB. If it only talks to remote APIs, a 2 GB plan is enough. Running models locally with Ollama is what needs memory — 4 GB for small 3B models, 6 GB or more for 7–8B models.

Can I run local models on a CPU-only VPS?

Yes, but keep expectations realistic. Small quantized models answer at a few tokens per second on CPU. That's fine for short tasks and private experiments, not for a snappy chat with long answers.

Can I use it on a NAT plan?

Yes, through an SSH tunnel on your personal SSH port. For a public HTTPS address that your team can open in a browser, you need a dedicated-IP plan.

Who can create accounts?

The first account you register becomes the admin. After that, set new sign-ups to pending or disable them in the admin settings, or anyone who finds the URL can create an account.

Where are my chats stored?

In a Docker volume on your server. Nothing goes to a third party except the prompts you send to whichever model API you connect.

Comments

No comments yet. Be the first.

Leave a comment

Comments are moderated before they appear.