"The server is slow" is a useless bug report — unless you can open a graph and see that memory climbed for three days before the crash. That's what Prometheus and Grafana give you: every server's CPU, RAM, disk and network, recorded every 15 seconds and drawn as charts you can scroll back through. Put them on their own small VPS and they keep watching even when a production box falls over.
The pieces
- node_exporter on every server you want to watch — exposes system metrics on port 9100.
- Prometheus on the monitoring VPS — pulls (scrapes) those metrics and stores them.
- Grafana — the dashboards, behind HTTPS.
- Optionally Alertmanager or Grafana alerting to message you when something crosses a line.
What you need
- Micro-IP ($10/mo, 2 vCPU, 2 GB RAM, 25 GB) watches a dozen or two servers. More hosts, custom app metrics or months of retention: Small-IP.
- A dedicated IPv4 for the monitoring server — Grafana is served over HTTPS, and your targets can whitelist one stable address for port 9100.
1. node_exporter on each target
apt install -y prometheus-node-exporter
# allow only the monitoring server to scrape it
ufw allow from 203.0.113.10 to any port 9100 proto tcp
Replace 203.0.113.10 with your monitoring VPS address. Never leave 9100 open to the world — it leaks a lot about your server.
2. Prometheus and Grafana on the monitoring VPS
apt update && apt install -y docker.io docker-compose-v2 caddy
mkdir -p /opt/monitoring && cd /opt/monitoring
cat > prometheus.yml <<'EOF'
global:
scrape_interval: 15s
scrape_configs:
- job_name: nodes
static_configs:
- targets: ['198.51.100.21:9100', '198.51.100.22:9100']
EOF
cat > compose.yml <<'EOF'
services:
prometheus:
image: prom/prometheus
volumes: [ "./prometheus.yml:/etc/prometheus/prometheus.yml:ro", "prom:/prometheus" ]
command: [ "--config.file=/etc/prometheus/prometheus.yml", "--storage.tsdb.retention.time=15d", "--web.enable-remote-write-receiver" ]
ports: [ "127.0.0.1:9090:9090" ]
restart: unless-stopped
grafana:
image: grafana/grafana
volumes: [ "grafana:/var/lib/grafana" ]
ports: [ "127.0.0.1:3000:3000" ]
restart: unless-stopped
volumes: { prom: {}, grafana: {} }
EOF
docker compose up -d
HTTPS in front of Grafana:
echo 'grafana.example.com {
reverse_proxy 127.0.0.1:3000
}' > /etc/caddy/Caddyfile && systemctl reload caddy
Log in (admin / admin, then change it), add Prometheus as a data source at http://prometheus:9090, and import the community "Node Exporter Full" dashboard (ID 1860). You get a full system dashboard in about a minute.
Servers on NAT plans
A NAT server only accepts inbound connections on its personal SSH port, so Prometheus can't scrape its exporter. Flip the direction: run Prometheus in agent mode on the NAT box and push to the monitoring server, which already accepts remote writes (the --web.enable-remote-write-receiver flag above).
# /etc/prometheus/agent.yml on the NAT server
scrape_configs:
- job_name: self
static_configs: [ { targets: ['127.0.0.1:9100'] } ]
remote_write:
- url: https://prom-push.example.com/api/v1/write
Expose that push endpoint through Caddy with basic auth rather than opening port 9090 raw.
Honest notes
Prometheus pulls; that's great for a fleet you control and awkward for anything behind NAT — hence the agent trick. Its local storage is not meant for years of history; for long retention you'd add a remote store later. And dashboards are only half the value: set at least three alerts on day one (disk above 85%, memory pressure, target down), or you'll only look at the graphs after something has already broken.
Related
- VPS for uptime monitoring (Uptime Kuma) — outage alerts and status pages
- How to configure a UFW firewall
- How much RAM does a VPS need?
Comments
No comments yet. Be the first.