Skip to content

Troubleshooting Methodology

When facing issues, blindly searching and trial-and-error often wastes a great deal of time. This article introduces a general troubleshooting framework suitable for EL systems.

  1. Observe first, act later — Gather sufficient information before modifying any configuration
  2. Start close, then go far — Check the most recent changes first, then expand the investigation scope
  3. Change one variable at a time — Modify only one thing at a time, verify, then move to the next
  4. Document the process — Record what you did and what you observed
  1. Clearly define the symptoms

    Can you describe the problem specifically? “The website is down” is not specific enough — “Accessing port 80 returns connection refused” is an effective description.

  2. Check recent changes

    View recent system logs
    $ sudo journalctl --since "1 hour ago" --priority err
    View recently installed/updated packages
    $ dnf history list --reverse | tail -20
  3. Check service status

    View the status and logs of the relevant service
    $ sudo systemctl status <service-name>
    $ sudo journalctl -u <service-name> -n 50 --no-pager
  4. Check resources

    Disk space
    $ df -h
    Memory usage
    $ free -h
    CPU and processes
    $ top -bn1 | head -20
  5. Check networking

    Port listening status
    $ sudo ss -tlnp
    Firewall rules
    $ sudo firewall-cmd --list-all
  6. Check SELinux

    SELinux denials are a very common source of issues on EL systems:

    View SELinux denial logs
    $ sudo ausearch -m avc -ts recent
    Check current SELinux mode
    $ getenforce
  7. Review log files

    System logs
    $ sudo journalctl -xe
    Application-specific logs (nginx as an example)
    $ sudo tail -50 /var/log/nginx/error.log
TargetCommand
All system logsjournalctl
Current boot logsjournalctl -b
Specific service logsjournalctl -u sshd
Follow logs in real timejournalctl -f
Error level and abovejournalctl -p err
Specific time rangejournalctl --since "2024-01-01" --until "2024-01-02"
  • Do not disable SELinux as a first resort — First investigate whether it is an SELinux policy issue
  • Do not blindly chmod 777 — This introduces security risks; find the correct permission settings instead
  • Do not overlook disk space — A full /var/log is an extremely common cause of failures
  • Do not forget the firewall — If a service is running but inaccessible externally, it is very likely a firewalld rule issue