High Availability Basics
The goal of High Availability (HA) is to eliminate single points of failure, ensuring that services continue to operate even when some components fail. This article introduces the core concepts of HA and commonly used implementation approaches on RHEL-based distributions.
Core Concepts
Section titled “Core Concepts”High Availability (HA)
Section titled “High Availability (HA)”High availability refers to the ability of a system to provide service within an agreed-upon timeframe. It is typically measured in “nines”:
| Availability | Allowed Downtime per Year |
|---|---|
| 99.9% (three nines) | 8 hours 45 minutes |
| 99.99% (four nines) | 52 minutes |
| 99.999% (five nines) | 5 minutes |
Failover
Section titled “Failover”The process by which a standby node automatically takes over when the primary node fails. There are two types:
- Active-Passive: Only the primary node serves traffic under normal conditions; the backup node stands by
- Active-Active: All nodes serve traffic simultaneously; when one fails, traffic is automatically distributed to the remaining nodes
Load Balancing
Section titled “Load Balancing”Distributes traffic across multiple backend servers, increasing both processing capacity and redundancy.
Keepalived — VRRP Virtual IP Failover
Section titled “Keepalived — VRRP Virtual IP Failover”Keepalived uses VRRP (Virtual Router Redundancy Protocol) to share a Virtual IP (VIP) among multiple servers. When the primary node fails, the VIP automatically floats to the backup node.
Installation
Section titled “Installation”Run the following on both primary and backup servers:
sudo dnf install keepalived -yConfiguring the Primary Node
Section titled “Configuring the Primary Node”sudo tee /etc/keepalived/keepalived.conf > /dev/null <<'EOF'global_defs { router_id LVS_MASTER script_user root enable_script_security}
# Health check scriptvrrp_script check_nginx { script "/usr/local/bin/check_nginx.sh" interval 2 weight -20 fall 3 rise 2}
vrrp_instance VI_1 { state MASTER interface eth0 # Adjust to your actual NIC name virtual_router_id 51 priority 100 # Higher priority for the primary node advert_int 1
authentication { auth_type PASS auth_pass mySecretPass # Must be the same on primary and backup }
virtual_ipaddress { 192.168.1.100/24 # Virtual IP address }
track_script { check_nginx }
# Notification scripts triggered on VIP failover (optional) notify_master "/usr/local/bin/notify.sh MASTER" notify_backup "/usr/local/bin/notify.sh BACKUP" notify_fault "/usr/local/bin/notify.sh FAULT"}EOFConfiguring the Backup Node
Section titled “Configuring the Backup Node”sudo tee /etc/keepalived/keepalived.conf > /dev/null <<'EOF'global_defs { router_id LVS_BACKUP script_user root enable_script_security}
vrrp_script check_nginx { script "/usr/local/bin/check_nginx.sh" interval 2 weight -20 fall 3 rise 2}
vrrp_instance VI_1 { state BACKUP interface eth0 virtual_router_id 51 priority 90 # Lower priority for the backup node advert_int 1
authentication { auth_type PASS auth_pass mySecretPass }
virtual_ipaddress { 192.168.1.100/24 }
track_script { check_nginx }
notify_master "/usr/local/bin/notify.sh MASTER" notify_backup "/usr/local/bin/notify.sh BACKUP" notify_fault "/usr/local/bin/notify.sh FAULT"}EOFHealth Check Script
Section titled “Health Check Script”sudo tee /usr/local/bin/check_nginx.sh > /dev/null <<'EOF'#!/bin/bash# Check if nginx is running properlyif ! pidof nginx > /dev/null 2>&1; then # Attempt restart systemctl start nginx sleep 2 if ! pidof nginx > /dev/null 2>&1; then exit 1 # Non-zero return indicates check failure fifiexit 0EOF
chmod +x /usr/local/bin/check_nginx.shNotification Script
Section titled “Notification Script”sudo tee /usr/local/bin/notify.sh > /dev/null <<'EOF'#!/bin/bashSTATE=$1HOSTNAME=$(hostname)DATE=$(date '+%Y-%m-%d %H:%M:%S')MAILTO="[email protected]"
echo "[${DATE}] ${HOSTNAME} transitioned to ${STATE}" >> /var/log/keepalived_notify.logecho "${HOSTNAME} transitioned to ${STATE} at ${DATE}" | \ mail -s "[HA Alert] ${HOSTNAME} -> ${STATE}" "$MAILTO"EOF
chmod +x /usr/local/bin/notify.shStarting the Service
Section titled “Starting the Service”Run the following on both primary and backup nodes:
sudo systemctl enable --now keepalived
# Allow the VRRP protocol through the firewallsudo firewall-cmd --permanent --add-rich-rule='rule protocol value="vrrp" accept'sudo firewall-cmd --reloadVerification
Section titled “Verification”# Check VIP bindingip addr show eth0
# On the primary node, you should see the virtual IP# 2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> ...# inet 192.168.1.10/24 ...# inet 192.168.1.100/24 scope global secondary eth0
# Simulate a failure: stop keepalived on the primary nodesudo systemctl stop keepalived
# Verify on the backup node that the VIP has floated overip addr show eth0HAProxy — Load Balancing
Section titled “HAProxy — Load Balancing”HAProxy is a high-performance TCP/HTTP load balancer, commonly used in combination with Keepalived.
Installation
Section titled “Installation”sudo dnf install haproxy -yBasic Configuration
Section titled “Basic Configuration”sudo cp /etc/haproxy/haproxy.cfg /etc/haproxy/haproxy.cfg.bak
sudo tee /etc/haproxy/haproxy.cfg > /dev/null <<'EOF'global log 127.0.0.1 local2 chroot /var/lib/haproxy pidfile /var/run/haproxy.pid maxconn 4000 user haproxy group haproxy daemon
# SSL settings ssl-default-bind-ciphers PROFILE=SYSTEM ssl-default-server-ciphers PROFILE=SYSTEM
defaults mode http log global option httplog option dontlognull option http-server-close option forwardfor except 127.0.0.0/8 retries 3 timeout http-request 10s timeout queue 1m timeout connect 10s timeout client 1m timeout server 1m timeout http-keep-alive 10s timeout check 10s maxconn 3000
# Statistics pagelisten stats bind *:8404 stats enable stats uri /stats stats refresh 10s stats admin if LOCALHOST stats auth admin:YourStatsPassword
# HTTP frontendfrontend http_front bind *:80 default_backend web_servers
# HTTPS frontend (optional)frontend https_front bind *:443 ssl crt /etc/haproxy/certs/site.pem default_backend web_servers
# Backend server poolbackend web_servers balance roundrobin option httpchk GET /health http-check expect status 200
server web1 192.168.1.11:80 check inter 5s fall 3 rise 2 server web2 192.168.1.12:80 check inter 5s fall 3 rise 2 server web3 192.168.1.13:80 check inter 5s fall 3 rise 2 backupEOFCommon Load Balancing Algorithms
Section titled “Common Load Balancing Algorithms”balance roundrobin # Round robin (default), distributes requests sequentiallybalance leastconn # Least connections, prefers backends with fewest connectionsbalance source # Source IP hash, same client always reaches the same backendbalance uri # URI hash, same URI reaches the same backend (good for caching)Starting and Verifying
Section titled “Starting and Verifying”# Check configuration file syntaxhaproxy -c -f /etc/haproxy/haproxy.cfg
sudo systemctl enable --now haproxy
# Allow through the firewallsudo firewall-cmd --permanent --add-service=httpsudo firewall-cmd --permanent --add-service=httpssudo firewall-cmd --permanent --add-port=8404/tcpsudo firewall-cmd --reloadVisit http://server-ip:8404/stats to view the HAProxy statistics page.
HAProxy + Keepalived Architecture
Section titled “HAProxy + Keepalived Architecture”A typical high-availability load balancing architecture:
VIP: 192.168.1.100 +------------------+ | Keepalived | | (VRRP failover) | +--------+---------+ +-------------+-------------+ +------+------+ +------+------+ | HAProxy 1 | | HAProxy 2 | | (MASTER) | | (BACKUP) | +------+------+ +------+------+ | | +----------+----------+ +----------+----------+ | | | | | | +--+--+ +--+--+ +--+--+ |Web 1| |Web 2| |Web 3| +-----+ +-----+ +-----+Two HAProxy servers share the VIP via Keepalived, and clients always connect to the VIP. When the primary HAProxy fails, the VIP floats to the backup node, providing seamless failover.
Pacemaker + Corosync
Section titled “Pacemaker + Corosync”Pacemaker/Corosync is an enterprise-grade HA cluster solution, suitable for managing complex multi-resource clusters.
- Corosync: Provides cluster communication and membership management
- Pacemaker: The cluster resource manager that decides which node runs each resource
Installation
Section titled “Installation”Run the following on all cluster nodes:
sudo dnf install pacemaker corosync pcs -y
# pcs is the management tool for Pacemaker/CorosyncInitializing the Cluster
Section titled “Initializing the Cluster”# Set the hacluster user password on all nodes (must be the same)sudo passwd hacluster
# Start the pcs daemonsudo systemctl enable --now pcsd
# Allow through the firewallsudo firewall-cmd --permanent --add-service=high-availabilitysudo firewall-cmd --reloadOn one of the nodes:
# Authenticate all nodessudo pcs host auth node1 node2 node3 -u hacluster -p YourPassword
# Create the clustersudo pcs cluster setup mycluster node1 node2 node3
# Start the clustersudo pcs cluster start --allsudo pcs cluster enable --allConfiguring Cluster Resources
Section titled “Configuring Cluster Resources”# View cluster statussudo pcs status
# Disable STONITH (test environment only; production must configure STONITH devices)sudo pcs property set stonith-enabled=false
# Set the policy for when the cluster loses quorumsudo pcs property set no-quorum-policy=ignore
# Add a Virtual IP resourcesudo pcs resource create VirtualIP ocf:heartbeat:IPaddr2 \ ip=192.168.1.100 \ cidr_netmask=24 \ op monitor interval=30s
# Add an Nginx service resourcesudo pcs resource create WebServer systemd:nginx \ op monitor interval=10s timeout=20s
# Ensure VIP and WebServer run on the same nodesudo pcs constraint colocation add WebServer with VirtualIP INFINITY
# Ensure VIP starts before WebServersudo pcs constraint order VirtualIP then WebServer
# Set a preference for the resource to run on node1sudo pcs constraint location WebServer prefers node1=100Common Management Commands
Section titled “Common Management Commands”# View detailed cluster statussudo pcs status --full
# View resource configurationsudo pcs resource config
# View constraintssudo pcs constraint show
# Manually move a resource to another nodesudo pcs resource move WebServer node2
# Clear migration constraints (restore automatic scheduling)sudo pcs resource clear WebServer
# Put a node in standby modesudo pcs node standby node1
# Restore a node from standbysudo pcs node unstandby node1
# Clear resource error statesudo pcs resource cleanup WebServerShared Storage Concepts
Section titled “Shared Storage Concepts”In HA clusters, multiple nodes may need to access the same data, which requires shared storage.
Common Solutions
Section titled “Common Solutions”| Solution | Description | Use Case |
|---|---|---|
| NFS | Network File System | Simple sharing, non-high-I/O scenarios |
| GFS2 | Global File System | Multi-node concurrent read/write |
| DRBD | Distributed Replicated Block Device | Active-passive data synchronization |
| Ceph | Distributed Storage System | Large-scale clusters requiring high reliability |
| iSCSI | Network Block Storage | SAN storage connectivity |
NFS Shared Storage Example
Section titled “NFS Shared Storage Example”NFS server side:
sudo dnf install nfs-utils -y
# Create the shared directorysudo mkdir -p /data/sharedsudo chown nobody:nobody /data/shared
# Configure exportssudo tee /etc/exports > /dev/null <<'EOF'/data/shared 192.168.1.0/24(rw,sync,no_subtree_check,no_root_squash)EOF
sudo systemctl enable --now nfs-serversudo exportfs -rav
# Firewallsudo firewall-cmd --permanent --add-service=nfssudo firewall-cmd --permanent --add-service=mountdsudo firewall-cmd --permanent --add-service=rpc-bindsudo firewall-cmd --reloadNFS client (cluster nodes):
sudo dnf install nfs-utils -y
# Mountsudo mkdir -p /mnt/sharedsudo mount -t nfs nfs-server:/data/shared /mnt/shared
# Add to fstab for automatic mounting at bootecho 'nfs-server:/data/shared /mnt/shared nfs defaults,_netdev 0 0' | \ sudo tee -a /etc/fstabDRBD Overview
Section titled “DRBD Overview”DRBD synchronizes block device data in real time between two servers, essentially acting as network RAID 1:
# Install DRBD (requires the ELRepo repository)sudo dnf install https://www.elrepo.org/elrepo-release-9.el9.elrepo.noarch.rpm -ysudo dnf install drbd90-utils kmod-drbd90 -y
# Load the kernel modulesudo modprobe drbdDetailed DRBD configuration is beyond the scope of this article, but its basic operating mode is as follows:
+----------+ +----------+ | Node 1 | Network | Node 2 | | (Primary) | <-------> |(Secondary)| | /dev/drbd0| |/dev/drbd0| +-----+----+ +-----+----+ | | /dev/sdb1 /dev/sdb1 (local disk) (local disk)Choosing an HA Solution
Section titled “Choosing an HA Solution”| Requirement | Recommended Solution |
|---|---|
| Web server redundancy | Keepalived + Nginx/HAProxy |
| Database high availability | Database-native replication + Keepalived |
| Complex multi-resource clusters | Pacemaker + Corosync |
| Large-scale load balancing | HAProxy + Keepalived |
| Data synchronization | DRBD (two nodes) / Ceph (multi-node) |