Most people guess their RAM. They either buy far too much "just in case" or discover the OOM killer at 3 a.m. Memory is actually one of the easiest things to size, because workloads are predictable once you know the rough numbers. Here they are, from a 1 GB bot box up to the 64 GB servers people search for when they want to run a large model locally.
Rough numbers by workload
These are real-world steady-state figures for a typical setup, not vendor minimums.
| Workload | RAM it actually uses | Plan that fits |
|---|---|---|
| Telegram/Discord bot (Python, Node.js) | 100–300 MB | Nano 1 GB |
| Static site + Caddy/Nginx | 50–150 MB | Nano 1 GB |
| WordPress + MariaDB, normal traffic | 0.8–1.5 GB | Micro-IP 2 GB |
| PostgreSQL for a small app | 0.5–2 GB (you decide via shared_buffers) | Micro / Small |
| n8n with a few workflows | 0.5–1.5 GB | Micro 2 GB |
| Docker host with 5–10 small services | 2–4 GB | Small 4 GB |
| Coolify + builds + a database | 3–4 GB | Small-IP 4 GB |
| 7–8B LLM, 4-bit, CPU inference | 5–6 GB | Medium 6 GB |
| 14B LLM, 4-bit | ~10 GB | Pro-32 |
| 32B LLM, 4-bit | ~20 GB | Pro-32 |
| 70B LLM, 4-bit | ~40–45 GB | Pro-48 / Pro-64 |
Two things to notice. First, the jump between "normal" workloads and LLMs is huge — a bot and a 70B model differ by a factor of two hundred. Second, most services are fine on 1–4 GB; people overbuy because they size for a future that rarely arrives.
The LLM formula
For language models there's a simple estimate: parameters × bits per weight ÷ 8, plus overhead. An 8B model at roughly 4.5 bits (a typical Q4 quantization) is 8 × 4.5 ÷ 8 ≈ 4.5 GB of weights. Add 1–2 GB for the context window (the KV cache grows with context length) and the runtime, and you land at 5–6 GB.
The honest part: on CPU-only servers memory is only half the story. Expect single-digit tokens per second for 7–8B models on a few vCPUs and roughly one token per second for 70B. That's fine for batch jobs, background agents and private experiments; it is not a chat experience for many simultaneous users. If you mostly call hosted model APIs, your agent needs far less — see VPS sizing for AI agents.
Measure instead of guessing
If you already have a server, the numbers are right there:
free -h # look at "available", not "free"
ps aux --sort=-rss | head -n 8 # the biggest processes, by resident memory
docker stats --no-stream # per-container usage
Linux uses idle RAM as disk cache, so "free" is always small and that's healthy. Available is the number that matters — memory the kernel can hand to programs right now. If available stays above ~25% during your busiest hour, you're sized right. If it regularly hits zero and swap grows, move up.
Check for past OOM kills too:
journalctl -k | grep -i "out of memory"
Swap: a seatbelt, not an engine
A 1–2 GB swap file on a small server is cheap insurance: a package upgrade or a Docker build that briefly needs more memory survives instead of getting killed. But if a service lives in swap, every request waits on disk. Swap is for spikes. Constant swapping means you need the next plan.
So how much should you buy?
- 1 GB — one bot, one small API, a static site.
- 2 GB — the sensible default for anything with a database or Docker.
- 4–6 GB — several services on one box, builds on the server, a small local model.
- 32–64 GB — only when one big thing must live in memory: a local model above 14B, a large RAG index, an agent fleet. That's what the high-memory plans are for.
Start one size smaller than your gut says, watch free -h for a week, and upgrade if the numbers say so. On EQVPS a plan upgrade keeps your data and charges only the pro-rated difference (how plan changes work).
Comments
No comments yet. Be the first.