System Monitoring
Good system monitoring is the cornerstone of operations work. This article covers a complete range of solutions from built-in tools to professional monitoring stacks, helping you detect hidden issues before they become problems and quickly identify root causes when failures occur.
Cockpit — Built-in Web Management Interface
Section titled “Cockpit — Built-in Web Management Interface”Cockpit is a lightweight web management console that ships with RHEL-based distributions. It works out of the box and is ideal for day-to-day monitoring and management of small to medium-scale servers.
Installation and Startup
Section titled “Installation and Startup”# AlmaLinux/Rocky Linux 9 comes with cockpit pre-installed; if not installed:sudo dnf install cockpit -y
# Start and enable at bootsudo systemctl enable --now cockpit.socket
# Allow through the firewall (Cockpit uses port 9090)sudo firewall-cmd --permanent --add-service=cockpitsudo firewall-cmd --reloadOnce installed, open https://your-server-ip:9090 in a browser and log in with your system account.
Common Extension Modules
Section titled “Common Extension Modules”# Install additional management modulessudo dnf install cockpit-storaged # Storage managementsudo dnf install cockpit-networkmanager # Network managementsudo dnf install cockpit-podman # Container managementsudo dnf install cockpit-machines # Virtual machine management (requires libvirt)Multi-Server Centralized Management
Section titled “Multi-Server Centralized Management”Cockpit supports managing multiple servers from a single host. Simply add remote hosts in the Cockpit interface on the control machine. Remote hosts also need Cockpit installed with port 9090 open.
# Run the same installation steps on the managed remote serverssudo dnf install cockpit -ysudo systemctl enable --now cockpit.socketsudo firewall-cmd --permanent --add-service=cockpitsudo firewall-cmd --reloadPrometheus + node_exporter
Section titled “Prometheus + node_exporter”For scenarios that require long-term data storage, flexible querying, and alerting, Prometheus is the de facto standard.
Installing node_exporter
Section titled “Installing node_exporter”node_exporter runs on each monitored server to collect system metrics.
# Create a dedicated usersudo useradd --no-create-home --shell /sbin/nologin node_exporter
# Download the latest version (adjust version number as needed)cd /tmpcurl -LO https://github.com/prometheus/node_exporter/releases/download/v1.8.2/node_exporter-1.8.2.linux-amd64.tar.gztar xzf node_exporter-1.8.2.linux-amd64.tar.gzsudo cp node_exporter-1.8.2.linux-amd64/node_exporter /usr/local/bin/sudo chown node_exporter:node_exporter /usr/local/bin/node_exporterCreate a systemd service file:
sudo tee /etc/systemd/system/node_exporter.service > /dev/null <<'EOF'[Unit]Description=Prometheus Node ExporterAfter=network-online.targetWants=network-online.target
[Service]User=node_exporterGroup=node_exporterType=simpleExecStart=/usr/local/bin/node_exporter \ --collector.systemd \ --collector.processes
[Install]WantedBy=multi-user.targetEOF
sudo systemctl daemon-reloadsudo systemctl enable --now node_exporter
# Allow through the firewallsudo firewall-cmd --permanent --add-port=9100/tcpsudo firewall-cmd --reloadVerify it is running:
curl http://localhost:9100/metrics | head -20Installing Prometheus Server
Section titled “Installing Prometheus Server”Install Prometheus on the monitoring host:
sudo useradd --no-create-home --shell /sbin/nologin prometheus
# Create directoriessudo mkdir -p /etc/prometheus /var/lib/prometheussudo chown prometheus:prometheus /var/lib/prometheus
# Download and installcd /tmpcurl -LO https://github.com/prometheus/prometheus/releases/download/v2.53.0/prometheus-2.53.0.linux-amd64.tar.gztar xzf prometheus-2.53.0.linux-amd64.tar.gzcd prometheus-2.53.0.linux-amd64
sudo cp prometheus promtool /usr/local/bin/sudo cp -r consoles console_libraries /etc/prometheus/sudo chown -R prometheus:prometheus /etc/prometheusWrite the configuration file:
sudo tee /etc/prometheus/prometheus.yml > /dev/null <<'EOF'global: scrape_interval: 15s evaluation_interval: 15s
scrape_configs: - job_name: "prometheus" static_configs: - targets: ["localhost:9090"]
- job_name: "nodes" static_configs: - targets: - "192.168.1.10:9100" - "192.168.1.11:9100" - "192.168.1.12:9100" labels: env: "production"EOF
sudo chown prometheus:prometheus /etc/prometheus/prometheus.ymlCreate the systemd service:
sudo tee /etc/systemd/system/prometheus.service > /dev/null <<'EOF'[Unit]Description=Prometheus MonitoringAfter=network-online.targetWants=network-online.target
[Service]User=prometheusGroup=prometheusType=simpleExecStart=/usr/local/bin/prometheus \ --config.file=/etc/prometheus/prometheus.yml \ --storage.tsdb.path=/var/lib/prometheus/ \ --storage.tsdb.retention.time=30d \ --web.console.templates=/etc/prometheus/consoles \ --web.console.libraries=/etc/prometheus/console_librariesExecReload=/bin/kill -HUP $MAINPIDRestart=on-failure
[Install]WantedBy=multi-user.targetEOF
sudo systemctl daemon-reloadsudo systemctl enable --now prometheus
sudo firewall-cmd --permanent --add-port=9090/tcpsudo firewall-cmd --reloadVisit http://monitoring-host-ip:9090 to open the Prometheus Web UI.
Grafana Visualization
Section titled “Grafana Visualization”Grafana provides beautiful dashboards for Prometheus data.
Installing Grafana
Section titled “Installing Grafana”# Add the official repositorysudo tee /etc/yum.repos.d/grafana.repo > /dev/null <<'EOF'[grafana]name=Grafana OSSbaseurl=https://rpm.grafana.comrepo_gpgcheck=1enabled=1gpgcheck=1gpgkey=https://rpm.grafana.com/gpg.keysslverify=1sslcacert=/etc/pki/tls/certs/ca-bundle.crtEOF
sudo dnf install grafana -ysudo systemctl enable --now grafana-server
# Allow through the firewallsudo firewall-cmd --permanent --add-port=3000/tcpsudo firewall-cmd --reloadConfiguring the Data Source
Section titled “Configuring the Data Source”Visit http://server-ip:3000 (default username and password are both admin), then:
- Go to Configuration > Data Sources > Add data source
- Select Prometheus
- Set the URL to
http://localhost:9090 - Click Save & Test
Importing Dashboards
Section titled “Importing Dashboards”We recommend using the community-maintained Node Exporter dashboard:
- Go to Dashboards > Import
- Enter dashboard ID
1860(Node Exporter Full) - Select the Prometheus data source you just configured
- Click Import
Basic Monitoring Scripts
Section titled “Basic Monitoring Scripts”When a full monitoring stack is not yet deployed, simple scripts can still be very useful.
System Load and Uptime Monitoring
Section titled “System Load and Uptime Monitoring”#!/bin/bash# Check system load, memory, and disk usage
HOSTNAME=$(hostname)DATE=$(date '+%Y-%m-%d %H:%M:%S')LOAD=$(uptime | awk -F'load average:' '{print $2}' | xargs)UPTIME=$(uptime -p)
# Number of CPU cores (used to determine if load is too high)CPU_CORES=$(nproc)
# Memory usageMEM_USAGE=$(free | awk '/Mem:/ {printf "%.1f%%", $3/$2 * 100}')
# Disk usage (check all mount points)DISK_ALERT=""while read -r line; do usage=$(echo "$line" | awk '{print $5}' | tr -d '%') mount=$(echo "$line" | awk '{print $6}') if [ "$usage" -gt 85 ]; then DISK_ALERT="${DISK_ALERT} [WARNING] ${mount} usage at ${usage}%\n" fidone < <(df -h --type=ext4 --type=xfs | tail -n +2)
echo "=============================="echo "System Status Report - ${DATE}"echo "Hostname: ${HOSTNAME}"echo "Uptime: ${UPTIME}"echo "CPU Cores: ${CPU_CORES}"echo "Load Average: ${LOAD}"echo "Memory Usage: ${MEM_USAGE}"echo "=============================="
if [ -n "$DISK_ALERT" ]; then echo -e "Disk Alerts:\n${DISK_ALERT}"fichmod +x /usr/local/bin/check_system.shDisk Space Monitoring and Alerting
Section titled “Disk Space Monitoring and Alerting”#!/bin/bash# Send alert email when disk usage exceeds the threshold
THRESHOLD=85HOSTNAME=$(hostname)
df -h --type=ext4 --type=xfs | tail -n +2 | while read -r line; do usage=$(echo "$line" | awk '{print $5}' | tr -d '%') partition=$(echo "$line" | awk '{print $1}') mount=$(echo "$line" | awk '{print $6}')
if [ "$usage" -ge "$THRESHOLD" ]; then SUBJECT="[Disk Alert] ${HOSTNAME}: ${mount} usage at ${usage}%" BODY="Host: ${HOSTNAME}\nPartition: ${partition}\nMount Point: ${mount}\nUsage: ${usage}%\nTime: $(date)" echo -e "$BODY" | mail -s "$SUBJECT" "$MAILTO" fidoneProcess Monitoring Script
Section titled “Process Monitoring Script”#!/bin/bash# Check if critical processes are running; attempt restart and alert if not
SERVICES=("nginx" "mysqld" "sshd")HOSTNAME=$(hostname)
for service in "${SERVICES[@]}"; do if ! systemctl is-active --quiet "$service"; then echo "[$(date)] ${service} is down, attempting restart..." >> /var/log/process_monitor.log
systemctl start "$service" sleep 2
if systemctl is-active --quiet "$service"; then STATUS="automatically recovered" else STATUS="restart failed, immediate attention required!" fi
SUBJECT="[Process Alert] ${HOSTNAME}: ${service} ${STATUS}" echo "${SUBJECT}" | mail -s "$SUBJECT" "$MAILTO" fidoneSetting Up Alerts
Section titled “Setting Up Alerts”Prometheus Alertmanager
Section titled “Prometheus Alertmanager”Install Alertmanager:
cd /tmpcurl -LO https://github.com/prometheus/alertmanager/releases/download/v0.27.0/alertmanager-0.27.0.linux-amd64.tar.gztar xzf alertmanager-0.27.0.linux-amd64.tar.gz
sudo cp alertmanager-0.27.0.linux-amd64/alertmanager /usr/local/bin/sudo cp alertmanager-0.27.0.linux-amd64/amtool /usr/local/bin/sudo mkdir -p /etc/alertmanagerConfigure email alerting:
sudo tee /etc/alertmanager/alertmanager.yml > /dev/null <<'EOF'global: smtp_smarthost: 'smtp.example.com:587' smtp_from: '[email protected]' smtp_auth_username: '[email protected]' smtp_auth_password: 'your_password' smtp_require_tls: true
route: group_by: ['alertname', 'instance'] group_wait: 30s group_interval: 5m repeat_interval: 4h receiver: 'email-admin'
receivers: - name: 'email-admin' email_configs: - to: '[email protected]' send_resolved: trueEOFCreate alert rules:
sudo tee /etc/prometheus/alert_rules.yml > /dev/null <<'EOF'groups: - name: System Alerts rules: # Instance down - alert: InstanceDown expr: up == 0 for: 2m labels: severity: critical annotations: summary: "Instance {{ $labels.instance }} is down" description: "{{ $labels.instance }} has been unresponsive for more than 2 minutes."
# High CPU usage - alert: HighCpuUsage expr: 100 - (avg by(instance) (irate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85 for: 5m labels: severity: warning annotations: summary: "{{ $labels.instance }} CPU usage is too high" description: "CPU usage has exceeded 85%, current value: {{ $value | printf \"%.1f\" }}%"
# High memory usage - alert: HighMemoryUsage expr: (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90 for: 5m labels: severity: warning annotations: summary: "{{ $labels.instance }} memory usage is too high" description: "Memory usage has exceeded 90%, current value: {{ $value | printf \"%.1f\" }}%"
# Low disk space - alert: DiskSpaceLow expr: (1 - node_filesystem_avail_bytes{fstype=~"ext4|xfs"} / node_filesystem_size_bytes) * 100 > 85 for: 5m labels: severity: warning annotations: summary: "{{ $labels.instance }} disk space is low" description: "{{ $labels.mountpoint }} usage exceeds 85%, current: {{ $value | printf \"%.1f\" }}%"EOFReference the alert rules and Alertmanager in the Prometheus configuration:
# Add to /etc/prometheus/prometheus.yml:sudo tee /etc/prometheus/prometheus.yml > /dev/null <<'EOF'global: scrape_interval: 15s evaluation_interval: 15s
alerting: alertmanagers: - static_configs: - targets: - localhost:9093
rule_files: - "alert_rules.yml"
scrape_configs: - job_name: "prometheus" static_configs: - targets: ["localhost:9090"]
- job_name: "nodes" static_configs: - targets: - "192.168.1.10:9100" - "192.168.1.11:9100"EOF
# Reload Prometheus configurationsudo systemctl restart prometheusUsing cron to Run Monitoring Scripts on a Schedule
Section titled “Using cron to Run Monitoring Scripts on a Schedule”# Edit crontabsudo crontab -e
# Check disk every 5 minutes*/5 * * * * /usr/local/bin/disk_alert.sh
# Check critical processes every 2 minutes*/2 * * * * /usr/local/bin/check_process.sh
# Send system status report every day at 8 AMCommon Real-Time Monitoring Commands
Section titled “Common Real-Time Monitoring Commands”In day-to-day operations, the following commands let you quickly assess system status:
# View system load and processes in real timetophtop # Requires installation: dnf install htop
# Check load averagesuptimecat /proc/loadavg
# Memory usage overviewfree -h
# Disk usage overviewdf -hTdu -sh /var/log/* # View size of each item in a directory
# I/O usageiostat -xz 1 5 # Requires installation: dnf install sysstatiotop # Requires installation: dnf install iotop
# Network connection statisticsss -tunlp # View listening portsss -s # Connection statistics summary
# View real-time network trafficiftop # Requires installation: dnf install iftopSummary
Section titled “Summary”| Scenario | Recommended Solution |
|---|---|
| Single server daily management | Cockpit |
| Quick checks | Shell scripts + cron |
| Long-term monitoring of multiple servers | Prometheus + node_exporter + Grafana |
| Alert notifications | Alertmanager / scripts + email |
Choose the right solution based on your server scale and team needs. For production environments, deploying at least Prometheus + Grafana is recommended for full observability.