Skip to content

System Monitoring

Good system monitoring is the cornerstone of operations work. This article covers a complete range of solutions from built-in tools to professional monitoring stacks, helping you detect hidden issues before they become problems and quickly identify root causes when failures occur.

Cockpit — Built-in Web Management Interface

Section titled “Cockpit — Built-in Web Management Interface”

Cockpit is a lightweight web management console that ships with RHEL-based distributions. It works out of the box and is ideal for day-to-day monitoring and management of small to medium-scale servers.

Terminal window
# AlmaLinux/Rocky Linux 9 comes with cockpit pre-installed; if not installed:
sudo dnf install cockpit -y
# Start and enable at boot
sudo systemctl enable --now cockpit.socket
# Allow through the firewall (Cockpit uses port 9090)
sudo firewall-cmd --permanent --add-service=cockpit
sudo firewall-cmd --reload

Once installed, open https://your-server-ip:9090 in a browser and log in with your system account.

Terminal window
# Install additional management modules
sudo dnf install cockpit-storaged # Storage management
sudo dnf install cockpit-networkmanager # Network management
sudo dnf install cockpit-podman # Container management
sudo dnf install cockpit-machines # Virtual machine management (requires libvirt)

Cockpit supports managing multiple servers from a single host. Simply add remote hosts in the Cockpit interface on the control machine. Remote hosts also need Cockpit installed with port 9090 open.

Terminal window
# Run the same installation steps on the managed remote servers
sudo dnf install cockpit -y
sudo systemctl enable --now cockpit.socket
sudo firewall-cmd --permanent --add-service=cockpit
sudo firewall-cmd --reload

For scenarios that require long-term data storage, flexible querying, and alerting, Prometheus is the de facto standard.

node_exporter runs on each monitored server to collect system metrics.

Terminal window
# Create a dedicated user
sudo useradd --no-create-home --shell /sbin/nologin node_exporter
# Download the latest version (adjust version number as needed)
cd /tmp
curl -LO https://github.com/prometheus/node_exporter/releases/download/v1.8.2/node_exporter-1.8.2.linux-amd64.tar.gz
tar xzf node_exporter-1.8.2.linux-amd64.tar.gz
sudo cp node_exporter-1.8.2.linux-amd64/node_exporter /usr/local/bin/
sudo chown node_exporter:node_exporter /usr/local/bin/node_exporter

Create a systemd service file:

Terminal window
sudo tee /etc/systemd/system/node_exporter.service > /dev/null <<'EOF'
[Unit]
Description=Prometheus Node Exporter
After=network-online.target
Wants=network-online.target
[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter \
--collector.systemd \
--collector.processes
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now node_exporter
# Allow through the firewall
sudo firewall-cmd --permanent --add-port=9100/tcp
sudo firewall-cmd --reload

Verify it is running:

Terminal window
curl http://localhost:9100/metrics | head -20

Install Prometheus on the monitoring host:

Terminal window
sudo useradd --no-create-home --shell /sbin/nologin prometheus
# Create directories
sudo mkdir -p /etc/prometheus /var/lib/prometheus
sudo chown prometheus:prometheus /var/lib/prometheus
# Download and install
cd /tmp
curl -LO https://github.com/prometheus/prometheus/releases/download/v2.53.0/prometheus-2.53.0.linux-amd64.tar.gz
tar xzf prometheus-2.53.0.linux-amd64.tar.gz
cd prometheus-2.53.0.linux-amd64
sudo cp prometheus promtool /usr/local/bin/
sudo cp -r consoles console_libraries /etc/prometheus/
sudo chown -R prometheus:prometheus /etc/prometheus

Write the configuration file:

Terminal window
sudo tee /etc/prometheus/prometheus.yml > /dev/null <<'EOF'
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
- job_name: "prometheus"
static_configs:
- targets: ["localhost:9090"]
- job_name: "nodes"
static_configs:
- targets:
- "192.168.1.10:9100"
- "192.168.1.11:9100"
- "192.168.1.12:9100"
labels:
env: "production"
EOF
sudo chown prometheus:prometheus /etc/prometheus/prometheus.yml

Create the systemd service:

Terminal window
sudo tee /etc/systemd/system/prometheus.service > /dev/null <<'EOF'
[Unit]
Description=Prometheus Monitoring
After=network-online.target
Wants=network-online.target
[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
--config.file=/etc/prometheus/prometheus.yml \
--storage.tsdb.path=/var/lib/prometheus/ \
--storage.tsdb.retention.time=30d \
--web.console.templates=/etc/prometheus/consoles \
--web.console.libraries=/etc/prometheus/console_libraries
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now prometheus
sudo firewall-cmd --permanent --add-port=9090/tcp
sudo firewall-cmd --reload

Visit http://monitoring-host-ip:9090 to open the Prometheus Web UI.

Grafana provides beautiful dashboards for Prometheus data.

Terminal window
# Add the official repository
sudo tee /etc/yum.repos.d/grafana.repo > /dev/null <<'EOF'
[grafana]
name=Grafana OSS
baseurl=https://rpm.grafana.com
repo_gpgcheck=1
enabled=1
gpgcheck=1
gpgkey=https://rpm.grafana.com/gpg.key
sslverify=1
sslcacert=/etc/pki/tls/certs/ca-bundle.crt
EOF
sudo dnf install grafana -y
sudo systemctl enable --now grafana-server
# Allow through the firewall
sudo firewall-cmd --permanent --add-port=3000/tcp
sudo firewall-cmd --reload

Visit http://server-ip:3000 (default username and password are both admin), then:

  1. Go to Configuration > Data Sources > Add data source
  2. Select Prometheus
  3. Set the URL to http://localhost:9090
  4. Click Save & Test

We recommend using the community-maintained Node Exporter dashboard:

  1. Go to Dashboards > Import
  2. Enter dashboard ID 1860 (Node Exporter Full)
  3. Select the Prometheus data source you just configured
  4. Click Import

When a full monitoring stack is not yet deployed, simple scripts can still be very useful.

/usr/local/bin/check_system.sh
#!/bin/bash
# Check system load, memory, and disk usage
HOSTNAME=$(hostname)
DATE=$(date '+%Y-%m-%d %H:%M:%S')
LOAD=$(uptime | awk -F'load average:' '{print $2}' | xargs)
UPTIME=$(uptime -p)
# Number of CPU cores (used to determine if load is too high)
CPU_CORES=$(nproc)
# Memory usage
MEM_USAGE=$(free | awk '/Mem:/ {printf "%.1f%%", $3/$2 * 100}')
# Disk usage (check all mount points)
DISK_ALERT=""
while read -r line; do
usage=$(echo "$line" | awk '{print $5}' | tr -d '%')
mount=$(echo "$line" | awk '{print $6}')
if [ "$usage" -gt 85 ]; then
DISK_ALERT="${DISK_ALERT} [WARNING] ${mount} usage at ${usage}%\n"
fi
done < <(df -h --type=ext4 --type=xfs | tail -n +2)
echo "=============================="
echo "System Status Report - ${DATE}"
echo "Hostname: ${HOSTNAME}"
echo "Uptime: ${UPTIME}"
echo "CPU Cores: ${CPU_CORES}"
echo "Load Average: ${LOAD}"
echo "Memory Usage: ${MEM_USAGE}"
echo "=============================="
if [ -n "$DISK_ALERT" ]; then
echo -e "Disk Alerts:\n${DISK_ALERT}"
fi
Terminal window
chmod +x /usr/local/bin/check_system.sh
/usr/local/bin/disk_alert.sh
#!/bin/bash
# Send alert email when disk usage exceeds the threshold
THRESHOLD=85
HOSTNAME=$(hostname)
df -h --type=ext4 --type=xfs | tail -n +2 | while read -r line; do
usage=$(echo "$line" | awk '{print $5}' | tr -d '%')
partition=$(echo "$line" | awk '{print $1}')
mount=$(echo "$line" | awk '{print $6}')
if [ "$usage" -ge "$THRESHOLD" ]; then
SUBJECT="[Disk Alert] ${HOSTNAME}: ${mount} usage at ${usage}%"
BODY="Host: ${HOSTNAME}\nPartition: ${partition}\nMount Point: ${mount}\nUsage: ${usage}%\nTime: $(date)"
echo -e "$BODY" | mail -s "$SUBJECT" "$MAILTO"
fi
done
/usr/local/bin/check_process.sh
#!/bin/bash
# Check if critical processes are running; attempt restart and alert if not
SERVICES=("nginx" "mysqld" "sshd")
HOSTNAME=$(hostname)
for service in "${SERVICES[@]}"; do
if ! systemctl is-active --quiet "$service"; then
echo "[$(date)] ${service} is down, attempting restart..." >> /var/log/process_monitor.log
systemctl start "$service"
sleep 2
if systemctl is-active --quiet "$service"; then
STATUS="automatically recovered"
else
STATUS="restart failed, immediate attention required!"
fi
SUBJECT="[Process Alert] ${HOSTNAME}: ${service} ${STATUS}"
echo "${SUBJECT}" | mail -s "$SUBJECT" "$MAILTO"
fi
done

Install Alertmanager:

Terminal window
cd /tmp
curl -LO https://github.com/prometheus/alertmanager/releases/download/v0.27.0/alertmanager-0.27.0.linux-amd64.tar.gz
tar xzf alertmanager-0.27.0.linux-amd64.tar.gz
sudo cp alertmanager-0.27.0.linux-amd64/alertmanager /usr/local/bin/
sudo cp alertmanager-0.27.0.linux-amd64/amtool /usr/local/bin/
sudo mkdir -p /etc/alertmanager

Configure email alerting:

Terminal window
sudo tee /etc/alertmanager/alertmanager.yml > /dev/null <<'EOF'
global:
smtp_smarthost: 'smtp.example.com:587'
smtp_from: '[email protected]'
smtp_auth_username: '[email protected]'
smtp_auth_password: 'your_password'
smtp_require_tls: true
route:
group_by: ['alertname', 'instance']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receiver: 'email-admin'
receivers:
- name: 'email-admin'
email_configs:
send_resolved: true
EOF

Create alert rules:

Terminal window
sudo tee /etc/prometheus/alert_rules.yml > /dev/null <<'EOF'
groups:
- name: System Alerts
rules:
# Instance down
- alert: InstanceDown
expr: up == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Instance {{ $labels.instance }} is down"
description: "{{ $labels.instance }} has been unresponsive for more than 2 minutes."
# High CPU usage
- alert: HighCpuUsage
expr: 100 - (avg by(instance) (irate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }} CPU usage is too high"
description: "CPU usage has exceeded 85%, current value: {{ $value | printf \"%.1f\" }}%"
# High memory usage
- alert: HighMemoryUsage
expr: (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }} memory usage is too high"
description: "Memory usage has exceeded 90%, current value: {{ $value | printf \"%.1f\" }}%"
# Low disk space
- alert: DiskSpaceLow
expr: (1 - node_filesystem_avail_bytes{fstype=~"ext4|xfs"} / node_filesystem_size_bytes) * 100 > 85
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }} disk space is low"
description: "{{ $labels.mountpoint }} usage exceeds 85%, current: {{ $value | printf \"%.1f\" }}%"
EOF

Reference the alert rules and Alertmanager in the Prometheus configuration:

Terminal window
# Add to /etc/prometheus/prometheus.yml:
sudo tee /etc/prometheus/prometheus.yml > /dev/null <<'EOF'
global:
scrape_interval: 15s
evaluation_interval: 15s
alerting:
alertmanagers:
- static_configs:
- targets:
- localhost:9093
rule_files:
- "alert_rules.yml"
scrape_configs:
- job_name: "prometheus"
static_configs:
- targets: ["localhost:9090"]
- job_name: "nodes"
static_configs:
- targets:
- "192.168.1.10:9100"
- "192.168.1.11:9100"
EOF
# Reload Prometheus configuration
sudo systemctl restart prometheus

Using cron to Run Monitoring Scripts on a Schedule

Section titled “Using cron to Run Monitoring Scripts on a Schedule”
Terminal window
# Edit crontab
sudo crontab -e
# Check disk every 5 minutes
*/5 * * * * /usr/local/bin/disk_alert.sh
# Check critical processes every 2 minutes
*/2 * * * * /usr/local/bin/check_process.sh
# Send system status report every day at 8 AM
0 8 * * * /usr/local/bin/check_system.sh | mail -s "Daily System Report" [email protected]

In day-to-day operations, the following commands let you quickly assess system status:

Terminal window
# View system load and processes in real time
top
htop # Requires installation: dnf install htop
# Check load averages
uptime
cat /proc/loadavg
# Memory usage overview
free -h
# Disk usage overview
df -hT
du -sh /var/log/* # View size of each item in a directory
# I/O usage
iostat -xz 1 5 # Requires installation: dnf install sysstat
iotop # Requires installation: dnf install iotop
# Network connection statistics
ss -tunlp # View listening ports
ss -s # Connection statistics summary
# View real-time network traffic
iftop # Requires installation: dnf install iftop
ScenarioRecommended Solution
Single server daily managementCockpit
Quick checksShell scripts + cron
Long-term monitoring of multiple serversPrometheus + node_exporter + Grafana
Alert notificationsAlertmanager / scripts + email

Choose the right solution based on your server scale and team needs. For production environments, deploying at least Prometheus + Grafana is recommended for full observability.