−25%

on annual Windows plans, until 31 Oct. See plans

EQVPS

How much RAM does a VPS need? Real numbers by workload

Sep 26, 2026 · 4 min read · EQVPS Team

Most people guess their RAM. They either buy far too much "just in case" or discover the OOM killer at 3 a.m. Memory is actually one of the easiest things to size, because workloads are predictable once you know the rough numbers. Here they are, from a 1 GB bot box up to the 64 GB servers people search for when they want to run a large model locally.

Rough numbers by workload

These are real-world steady-state figures for a typical setup, not vendor minimums.

WorkloadRAM it actually usesPlan that fits
Telegram/Discord bot (Python, Node.js)100–300 MBNano 1 GB
Static site + Caddy/Nginx50–150 MBNano 1 GB
WordPress + MariaDB, normal traffic0.8–1.5 GBMicro-IP 2 GB
PostgreSQL for a small app0.5–2 GB (you decide via shared_buffers)Micro / Small
n8n with a few workflows0.5–1.5 GBMicro 2 GB
Docker host with 5–10 small services2–4 GBSmall 4 GB
Coolify + builds + a database3–4 GBSmall-IP 4 GB
7–8B LLM, 4-bit, CPU inference5–6 GBMedium 6 GB
14B LLM, 4-bit~10 GBPro-32
32B LLM, 4-bit~20 GBPro-32
70B LLM, 4-bit~40–45 GBPro-48 / Pro-64

Two things to notice. First, the jump between "normal" workloads and LLMs is huge — a bot and a 70B model differ by a factor of two hundred. Second, most services are fine on 1–4 GB; people overbuy because they size for a future that rarely arrives.

The LLM formula

For language models there's a simple estimate: parameters × bits per weight ÷ 8, plus overhead. An 8B model at roughly 4.5 bits (a typical Q4 quantization) is 8 × 4.5 ÷ 8 ≈ 4.5 GB of weights. Add 1–2 GB for the context window (the KV cache grows with context length) and the runtime, and you land at 5–6 GB.

The honest part: on CPU-only servers memory is only half the story. Expect single-digit tokens per second for 7–8B models on a few vCPUs and roughly one token per second for 70B. That's fine for batch jobs, background agents and private experiments; it is not a chat experience for many simultaneous users. If you mostly call hosted model APIs, your agent needs far less — see VPS sizing for AI agents.

Measure instead of guessing

If you already have a server, the numbers are right there:

free -h                              # look at "available", not "free"
ps aux --sort=-rss | head -n 8       # the biggest processes, by resident memory
docker stats --no-stream             # per-container usage

Linux uses idle RAM as disk cache, so "free" is always small and that's healthy. Available is the number that matters — memory the kernel can hand to programs right now. If available stays above ~25% during your busiest hour, you're sized right. If it regularly hits zero and swap grows, move up.

Check for past OOM kills too:

journalctl -k | grep -i "out of memory"

Swap: a seatbelt, not an engine

A 1–2 GB swap file on a small server is cheap insurance: a package upgrade or a Docker build that briefly needs more memory survives instead of getting killed. But if a service lives in swap, every request waits on disk. Swap is for spikes. Constant swapping means you need the next plan.

So how much should you buy?

Start one size smaller than your gut says, watch free -h for a week, and upgrade if the numbers say so. On EQVPS a plan upgrade keeps your data and charges only the pro-rated difference (how plan changes work).

FAQ

How much RAM does a basic VPS need?

For a single bot, a small API or a static site, 1 GB is enough — often with room to spare. 2 GB is the comfortable default once you add a database or run things in Docker. Below 1 GB, modern package managers and language runtimes start to hurt.

How much RAM do I need to run an LLM on a VPS?

Roughly the model size at its quantization plus a gigabyte or two for context. A 7–8B model at 4-bit needs about 5–6 GB, a 14B model about 10 GB, 32B around 20 GB and 70B around 40–45 GB. That's why 64 GB servers exist: they fit a 70B model with room for the rest of your stack.

Is swap a replacement for RAM?

No. Swap is a safety net for short spikes — it keeps a build or a backup from being killed. If a workload lives in swap all day, everything becomes slow; that's the signal to move up a plan.

How do I see how much RAM my server really uses?

Run 'free -h' and look at the 'available' column, not 'free' — Linux uses spare memory for disk cache and gives it back on demand. 'ps aux --sort=-rss | head' shows the biggest processes, and 'docker stats --no-stream' shows per-container usage.

When does a 64 GB VPS make sense?

When a single thing needs to sit in memory: a 32–70B local model, a large vector index for RAG, or a fleet of agents with their own runtimes. For websites, bots and most APIs it's wasted money.

← Back to blogSee plans & pricing →

Comments

No comments yet. Be the first.

Leave a comment

Comments are moderated before they appear.