Skip to content

Performance Issues

The first step in performance troubleshooting is identifying the bottleneck: CPU, memory, disk I/O, or network. This article introduces commonly used performance analysis tools and troubleshooting methods.

Before diving into analysis, get a snapshot of the overall system status.

At a glance: check load with uptime
$ uptime

Example output: 10:32:05 up 45 days, load average: 2.50, 1.80, 0.95

The load average values represent the 1-minute, 5-minute, and 15-minute averages. As a rule of thumb, load should not consistently exceed the number of CPU cores.

Check the number of CPU cores
$ nproc
Quick check of resource usage
$ free -h # Memory
$ df -h # Disk space
$ uptime # Load
$ sudo ss -s # Network connection statistics
Launch top
$ top

Key metrics in the top interface:

  • %us — User-space CPU usage
  • %sy — Kernel-space CPU usage
  • %wa — CPU time waiting for I/O (high values indicate disk is the bottleneck)
  • %id — Idle CPU time

Common operations:

  • Press 1 — Show per-core CPU usage
  • Press P — Sort by CPU usage
  • Press M — Sort by memory usage
  • Press c — Show full command line
Non-interactive mode: output a single snapshot
$ top -bn1 | head -30

htop provides a more intuitive interface and needs to be installed from EPEL:

Install htop
$ sudo dnf install -y epel-release
$ sudo dnf install -y htop
Launch htop
$ htop

In htop, press F6 to select sort criteria and F5 to toggle the tree view.

Install the sysstat package
$ sudo dnf install -y sysstat
Display per-core CPU usage every second, 5 times
$ mpstat -P ALL 1 5
Find the top 10 processes by CPU usage
$ ps aux --sort=-%cpu | head -11
View thread-level CPU usage for a specific process
$ top -H -p <PID>
  • Application bugs — Infinite loops, inefficient algorithms
  • Too many processes/threads — High context-switching overhead
  • Interrupt storms — Check /proc/interrupts
  • Compilation or compression tasks — CPU-intensive operations
Check context switch frequency
$ vmstat 1 5

The cs column shows context switches per interval. Tens of thousands per second or more may impact performance.

View memory usage
$ free -h

Understanding the output:

  • total — Total physical memory
  • used — Memory in use
  • free — Completely free memory
  • buff/cache — Memory used as buffers/cache (reclaimable)
  • available — Actually available memory (free + reclaimable cache)

Important: Linux aggressively uses free memory for disk caching, which is normal behavior. To determine whether memory is insufficient, look at the available column, not the free column.

Output every second, 10 times
$ vmstat 1 10

Key columns:

  • si / so — Pages swapped in/out. Consistently greater than 0 indicates memory shortage
  • free — Free memory (KB)
  • buff / cache — Buffer and cache memory
Process list sorted by memory usage
$ ps aux --sort=-%mem | head -11
View detailed memory information for a specific process
$ cat /proc/<PID>/status | grep -E '(VmRSS|VmSize|VmSwap)'
View swap usage overview
$ swapon --show
See which processes are using swap
$ for pid in /proc/[0-9]*; do
name=$(cat $pid/comm 2>/dev/null)
swap=$(grep VmSwap $pid/status 2>/dev/null | awk '{print $2}')
[ -n "$swap" ] && [ "$swap" -gt 0 ] && echo "$swap kB - $name (PID: $(basename $pid))"
done | sort -rn | head -10
  • Memory leaks — An application continuously consuming more memory
  • Unreleased cache — Certain applications holding large caches
  • OOM Killer — The system automatically kills processes when memory is exhausted
View OOM Killer history
$ dmesg | grep -i "out of memory"
$ journalctl --since "7 days ago" | grep -i "oom-kill"
Output disk I/O statistics every second, 5 times
$ iostat -x 1 5

Key metrics:

  • %util — Disk utilization. Consistently near 100% means the disk is saturated
  • await — Average I/O wait time (milliseconds). Normal is 5-20ms for HDD, 0.5-2ms for SSD
  • r/s, w/s — Read/write operations per second (IOPS)
  • rkB/s, wkB/s — Read/write throughput per second
View only a specific disk
$ iostat -x /dev/sda 1 5

Using iotop to Identify I/O-Heavy Processes

Section titled “Using iotop to Identify I/O-Heavy Processes”
Install iotop
$ sudo dnf install -y iotop
View real-time I/O usage
$ sudo iotop -o

The -o flag shows only processes with I/O activity.

View filesystem disk space
$ df -h
View inode usage
$ df -i

Note: Even if disk space appears to be available, inode exhaustion can also prevent new files from being created.

Find the largest directories
$ sudo du -sh /var/log/* | sort -rh | head -10
Find large files
$ sudo find / -xdev -type f -size +100M -exec ls -lh {} \; 2>/dev/null | sort -k5 -rh | head -10
  • Oversized log files — /var/log is full
  • Database I/O intensive — Unoptimized queries
  • Frequent swap reads/writes — Actually caused by memory shortage
  • Degraded RAID — Array performance degradation
View network interface traffic (every second, 5 times)
$ sar -n DEV 1 5

Key metrics:

  • rxpck/s, txpck/s — Packets received/sent per second
  • rxkB/s, txkB/s — Bytes received/sent per second
  • %ifutil — Network interface utilization
View connection statistics overview
$ ss -s
View TIME_WAIT connection count
$ ss -ant | awk '{print $1}' | sort | uniq -c | sort -rn

A large number of TIME_WAIT connections may indicate excessive short-lived connections. A large number of CLOSE_WAIT may indicate an application bug.

Install iperf3
$ sudo dnf install -y iperf3
Start iperf3 on the server side
$ iperf3 -s
Test bandwidth from the client side
$ iperf3 -c <server-IP>
  • Bandwidth saturation — Traffic exceeds NIC or link capacity
  • Packet loss — Poor network quality or buffer overflow
  • Too many connections — Exceeding system or application limits
  • DNS latency — DNS lookup on every request

sar (System Activity Reporter) can review historical performance data.

Ensure sysstat is installed and enabled
$ sudo dnf install -y sysstat
$ sudo systemctl enable --now sysstat
View today's CPU usage history
$ sar -u
View today's memory usage history
$ sar -r
View today's disk I/O history
$ sar -d
View data for a specific date (e.g., the 20th)
$ sar -u -f /var/log/sa/sa20
View a specific time range
$ sar -u -s 08:00:00 -e 12:00:00

Quick Performance Issue Identification Flow

Section titled “Quick Performance Issue Identification Flow”

Follow this sequence to quickly identify most performance issues:

Compare load with CPU core count
$ echo "Load: $(cat /proc/loadavg | awk '{print $1, $2, $3}'), CPU cores: $(nproc)"
Use vmstat for initial assessment
$ vmstat 1 5
  • High us + sy → CPU bottleneck
  • High wa → Disk I/O bottleneck
  • si/so consistently above 0 → Memory shortage
  • All of the above normal but system is slow → Likely a network or application-layer issue

Based on the results from step 2, use the tools in the corresponding section for in-depth analysis.

Use nice to lower process priority
$ nice -n 19 <command>
Use cpulimit to limit CPU usage of a running process
$ sudo dnf install -y cpulimit
$ sudo cpulimit -p <PID> -l 50 # Limit to 50% CPU
Create a swap file (emergency)
$ sudo fallocate -l 2G /swapfile
$ sudo chmod 600 /swapfile
$ sudo mkswap /swapfile
$ sudo swapon /swapfile
Clear system cache (emergency, not recommended as routine practice)
$ sudo sync && sudo sysctl vm.drop_caches=3
Clean dnf cache
$ sudo dnf clean all
Clean old journal logs
$ sudo journalctl --vacuum-size=200M
Remove old kernels (keep current and one spare)
$ sudo dnf remove --oldinstallonly --setopt installonly_limit=2 kernel

Too Many Network Connections: Tune Kernel Parameters

Section titled “Too Many Network Connections: Tune Kernel Parameters”
View current connection limits
$ sysctl net.core.somaxconn
$ cat /proc/sys/net/ipv4/tcp_max_tw_buckets
Temporarily increase connection limits
$ sudo sysctl -w net.core.somaxconn=65535
$ sudo sysctl -w net.ipv4.tcp_max_tw_buckets=65535