Skip to content

High Availability Basics

The goal of High Availability (HA) is to eliminate single points of failure, ensuring that services continue to operate even when some components fail. This article introduces the core concepts of HA and commonly used implementation approaches on RHEL-based distributions.

High availability refers to the ability of a system to provide service within an agreed-upon timeframe. It is typically measured in “nines”:

AvailabilityAllowed Downtime per Year
99.9% (three nines)8 hours 45 minutes
99.99% (four nines)52 minutes
99.999% (five nines)5 minutes

The process by which a standby node automatically takes over when the primary node fails. There are two types:

  • Active-Passive: Only the primary node serves traffic under normal conditions; the backup node stands by
  • Active-Active: All nodes serve traffic simultaneously; when one fails, traffic is automatically distributed to the remaining nodes

Distributes traffic across multiple backend servers, increasing both processing capacity and redundancy.

Keepalived uses VRRP (Virtual Router Redundancy Protocol) to share a Virtual IP (VIP) among multiple servers. When the primary node fails, the VIP automatically floats to the backup node.

Run the following on both primary and backup servers:

Terminal window
sudo dnf install keepalived -y
Terminal window
sudo tee /etc/keepalived/keepalived.conf > /dev/null <<'EOF'
global_defs {
router_id LVS_MASTER
script_user root
enable_script_security
}
# Health check script
vrrp_script check_nginx {
script "/usr/local/bin/check_nginx.sh"
interval 2
weight -20
fall 3
rise 2
}
vrrp_instance VI_1 {
state MASTER
interface eth0 # Adjust to your actual NIC name
virtual_router_id 51
priority 100 # Higher priority for the primary node
advert_int 1
authentication {
auth_type PASS
auth_pass mySecretPass # Must be the same on primary and backup
}
virtual_ipaddress {
192.168.1.100/24 # Virtual IP address
}
track_script {
check_nginx
}
# Notification scripts triggered on VIP failover (optional)
notify_master "/usr/local/bin/notify.sh MASTER"
notify_backup "/usr/local/bin/notify.sh BACKUP"
notify_fault "/usr/local/bin/notify.sh FAULT"
}
EOF
Terminal window
sudo tee /etc/keepalived/keepalived.conf > /dev/null <<'EOF'
global_defs {
router_id LVS_BACKUP
script_user root
enable_script_security
}
vrrp_script check_nginx {
script "/usr/local/bin/check_nginx.sh"
interval 2
weight -20
fall 3
rise 2
}
vrrp_instance VI_1 {
state BACKUP
interface eth0
virtual_router_id 51
priority 90 # Lower priority for the backup node
advert_int 1
authentication {
auth_type PASS
auth_pass mySecretPass
}
virtual_ipaddress {
192.168.1.100/24
}
track_script {
check_nginx
}
notify_master "/usr/local/bin/notify.sh MASTER"
notify_backup "/usr/local/bin/notify.sh BACKUP"
notify_fault "/usr/local/bin/notify.sh FAULT"
}
EOF
sudo tee /usr/local/bin/check_nginx.sh > /dev/null <<'EOF'
#!/bin/bash
# Check if nginx is running properly
if ! pidof nginx > /dev/null 2>&1; then
# Attempt restart
systemctl start nginx
sleep 2
if ! pidof nginx > /dev/null 2>&1; then
exit 1 # Non-zero return indicates check failure
fi
fi
exit 0
EOF
chmod +x /usr/local/bin/check_nginx.sh
sudo tee /usr/local/bin/notify.sh > /dev/null <<'EOF'
#!/bin/bash
STATE=$1
HOSTNAME=$(hostname)
DATE=$(date '+%Y-%m-%d %H:%M:%S')
echo "[${DATE}] ${HOSTNAME} transitioned to ${STATE}" >> /var/log/keepalived_notify.log
echo "${HOSTNAME} transitioned to ${STATE} at ${DATE}" | \
mail -s "[HA Alert] ${HOSTNAME} -> ${STATE}" "$MAILTO"
EOF
chmod +x /usr/local/bin/notify.sh

Run the following on both primary and backup nodes:

Terminal window
sudo systemctl enable --now keepalived
# Allow the VRRP protocol through the firewall
sudo firewall-cmd --permanent --add-rich-rule='rule protocol value="vrrp" accept'
sudo firewall-cmd --reload
Terminal window
# Check VIP binding
ip addr show eth0
# On the primary node, you should see the virtual IP
# 2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> ...
# inet 192.168.1.10/24 ...
# inet 192.168.1.100/24 scope global secondary eth0
# Simulate a failure: stop keepalived on the primary node
sudo systemctl stop keepalived
# Verify on the backup node that the VIP has floated over
ip addr show eth0

HAProxy is a high-performance TCP/HTTP load balancer, commonly used in combination with Keepalived.

Terminal window
sudo dnf install haproxy -y
Terminal window
sudo cp /etc/haproxy/haproxy.cfg /etc/haproxy/haproxy.cfg.bak
sudo tee /etc/haproxy/haproxy.cfg > /dev/null <<'EOF'
global
log 127.0.0.1 local2
chroot /var/lib/haproxy
pidfile /var/run/haproxy.pid
maxconn 4000
user haproxy
group haproxy
daemon
# SSL settings
ssl-default-bind-ciphers PROFILE=SYSTEM
ssl-default-server-ciphers PROFILE=SYSTEM
defaults
mode http
log global
option httplog
option dontlognull
option http-server-close
option forwardfor except 127.0.0.0/8
retries 3
timeout http-request 10s
timeout queue 1m
timeout connect 10s
timeout client 1m
timeout server 1m
timeout http-keep-alive 10s
timeout check 10s
maxconn 3000
# Statistics page
listen stats
bind *:8404
stats enable
stats uri /stats
stats refresh 10s
stats admin if LOCALHOST
stats auth admin:YourStatsPassword
# HTTP frontend
frontend http_front
bind *:80
default_backend web_servers
# HTTPS frontend (optional)
frontend https_front
bind *:443 ssl crt /etc/haproxy/certs/site.pem
default_backend web_servers
# Backend server pool
backend web_servers
balance roundrobin
option httpchk GET /health
http-check expect status 200
server web1 192.168.1.11:80 check inter 5s fall 3 rise 2
server web2 192.168.1.12:80 check inter 5s fall 3 rise 2
server web3 192.168.1.13:80 check inter 5s fall 3 rise 2 backup
EOF
balance roundrobin # Round robin (default), distributes requests sequentially
balance leastconn # Least connections, prefers backends with fewest connections
balance source # Source IP hash, same client always reaches the same backend
balance uri # URI hash, same URI reaches the same backend (good for caching)
Terminal window
# Check configuration file syntax
haproxy -c -f /etc/haproxy/haproxy.cfg
sudo systemctl enable --now haproxy
# Allow through the firewall
sudo firewall-cmd --permanent --add-service=http
sudo firewall-cmd --permanent --add-service=https
sudo firewall-cmd --permanent --add-port=8404/tcp
sudo firewall-cmd --reload

Visit http://server-ip:8404/stats to view the HAProxy statistics page.

A typical high-availability load balancing architecture:

VIP: 192.168.1.100
+------------------+
| Keepalived |
| (VRRP failover) |
+--------+---------+
+-------------+-------------+
+------+------+ +------+------+
| HAProxy 1 | | HAProxy 2 |
| (MASTER) | | (BACKUP) |
+------+------+ +------+------+
| |
+----------+----------+ +----------+----------+
| | | | | |
+--+--+ +--+--+ +--+--+
|Web 1| |Web 2| |Web 3|
+-----+ +-----+ +-----+

Two HAProxy servers share the VIP via Keepalived, and clients always connect to the VIP. When the primary HAProxy fails, the VIP floats to the backup node, providing seamless failover.

Pacemaker/Corosync is an enterprise-grade HA cluster solution, suitable for managing complex multi-resource clusters.

  • Corosync: Provides cluster communication and membership management
  • Pacemaker: The cluster resource manager that decides which node runs each resource

Run the following on all cluster nodes:

Terminal window
sudo dnf install pacemaker corosync pcs -y
# pcs is the management tool for Pacemaker/Corosync
Terminal window
# Set the hacluster user password on all nodes (must be the same)
sudo passwd hacluster
# Start the pcs daemon
sudo systemctl enable --now pcsd
# Allow through the firewall
sudo firewall-cmd --permanent --add-service=high-availability
sudo firewall-cmd --reload

On one of the nodes:

Terminal window
# Authenticate all nodes
sudo pcs host auth node1 node2 node3 -u hacluster -p YourPassword
# Create the cluster
sudo pcs cluster setup mycluster node1 node2 node3
# Start the cluster
sudo pcs cluster start --all
sudo pcs cluster enable --all
Terminal window
# View cluster status
sudo pcs status
# Disable STONITH (test environment only; production must configure STONITH devices)
sudo pcs property set stonith-enabled=false
# Set the policy for when the cluster loses quorum
sudo pcs property set no-quorum-policy=ignore
# Add a Virtual IP resource
sudo pcs resource create VirtualIP ocf:heartbeat:IPaddr2 \
ip=192.168.1.100 \
cidr_netmask=24 \
op monitor interval=30s
# Add an Nginx service resource
sudo pcs resource create WebServer systemd:nginx \
op monitor interval=10s timeout=20s
# Ensure VIP and WebServer run on the same node
sudo pcs constraint colocation add WebServer with VirtualIP INFINITY
# Ensure VIP starts before WebServer
sudo pcs constraint order VirtualIP then WebServer
# Set a preference for the resource to run on node1
sudo pcs constraint location WebServer prefers node1=100
Terminal window
# View detailed cluster status
sudo pcs status --full
# View resource configuration
sudo pcs resource config
# View constraints
sudo pcs constraint show
# Manually move a resource to another node
sudo pcs resource move WebServer node2
# Clear migration constraints (restore automatic scheduling)
sudo pcs resource clear WebServer
# Put a node in standby mode
sudo pcs node standby node1
# Restore a node from standby
sudo pcs node unstandby node1
# Clear resource error state
sudo pcs resource cleanup WebServer

In HA clusters, multiple nodes may need to access the same data, which requires shared storage.

SolutionDescriptionUse Case
NFSNetwork File SystemSimple sharing, non-high-I/O scenarios
GFS2Global File SystemMulti-node concurrent read/write
DRBDDistributed Replicated Block DeviceActive-passive data synchronization
CephDistributed Storage SystemLarge-scale clusters requiring high reliability
iSCSINetwork Block StorageSAN storage connectivity

NFS server side:

Terminal window
sudo dnf install nfs-utils -y
# Create the shared directory
sudo mkdir -p /data/shared
sudo chown nobody:nobody /data/shared
# Configure exports
sudo tee /etc/exports > /dev/null <<'EOF'
/data/shared 192.168.1.0/24(rw,sync,no_subtree_check,no_root_squash)
EOF
sudo systemctl enable --now nfs-server
sudo exportfs -rav
# Firewall
sudo firewall-cmd --permanent --add-service=nfs
sudo firewall-cmd --permanent --add-service=mountd
sudo firewall-cmd --permanent --add-service=rpc-bind
sudo firewall-cmd --reload

NFS client (cluster nodes):

Terminal window
sudo dnf install nfs-utils -y
# Mount
sudo mkdir -p /mnt/shared
sudo mount -t nfs nfs-server:/data/shared /mnt/shared
# Add to fstab for automatic mounting at boot
echo 'nfs-server:/data/shared /mnt/shared nfs defaults,_netdev 0 0' | \
sudo tee -a /etc/fstab

DRBD synchronizes block device data in real time between two servers, essentially acting as network RAID 1:

Terminal window
# Install DRBD (requires the ELRepo repository)
sudo dnf install https://www.elrepo.org/elrepo-release-9.el9.elrepo.noarch.rpm -y
sudo dnf install drbd90-utils kmod-drbd90 -y
# Load the kernel module
sudo modprobe drbd

Detailed DRBD configuration is beyond the scope of this article, but its basic operating mode is as follows:

+----------+ +----------+
| Node 1 | Network | Node 2 |
| (Primary) | <-------> |(Secondary)|
| /dev/drbd0| |/dev/drbd0|
+-----+----+ +-----+----+
| |
/dev/sdb1 /dev/sdb1
(local disk) (local disk)
RequirementRecommended Solution
Web server redundancyKeepalived + Nginx/HAProxy
Database high availabilityDatabase-native replication + Keepalived
Complex multi-resource clustersPacemaker + Corosync
Large-scale load balancingHAProxy + Keepalived
Data synchronizationDRBD (two nodes) / Ceph (multi-node)