NVIDIA Drivers and CUDA
Running AI workloads on EL (Enterprise Linux) systems requires proper installation of NVIDIA GPU drivers and the CUDA toolkit. This guide covers driver installation, CUDA configuration, nvidia-smi usage, and NVIDIA Container Toolkit deployment.
Check GPU Hardware
Section titled “Check GPU Hardware”Confirm the System Detects the GPU
Section titled “Confirm the System Detects the GPU”lspci | grep -i nvidiaExample output:
3b:00.0 3D controller: NVIDIA Corporation A100 80GB PCIe (rev a1)If there is no output, the system has not detected an NVIDIA GPU. Check the hardware connection or BIOS settings.
View Detailed GPU Information
Section titled “View Detailed GPU Information”lspci -v -s $(lspci | grep -i nvidia | head -1 | awk '{print $1}')Install NVIDIA Drivers
Section titled “Install NVIDIA Drivers”There are two main ways to install drivers: through the NVIDIA official repository (recommended) or through ELRepo.
Pre-Installation
Section titled “Pre-Installation”Ensure the system is updated and install necessary dependencies:
sudo dnf update -ysudo dnf install -y kernel-devel kernel-headers gcc make dkmsDisable the nouveau open-source driver (if loaded):
# Check if nouveau is loadedlsmod | grep nouveau
# If there is output, disable itsudo tee /etc/modprobe.d/blacklist-nouveau.conf > /dev/null <<'EOF'blacklist nouveauoptions nouveau modeset=0EOF
# Rebuild initramfssudo dracut --force
# Rebootsudo rebootOption 1: NVIDIA Official Repository (Recommended)
Section titled “Option 1: NVIDIA Official Repository (Recommended)”This is the recommended approach for getting the latest driver versions. Pick the repository that matches your EL release.
# Add the NVIDIA CUDA repository (includes both drivers and CUDA)sudo dnf config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel9/x86_64/cuda-rhel9.repo
# Clean cachesudo dnf clean all
# Install the latest driver (DKMS module)sudo dnf module install -y nvidia-driver:latest-dkms# Add the NVIDIA CUDA repository (rhel10)sudo dnf config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel10/x86_64/cuda-rhel10.repo
# Clean cachesudo dnf clean all
# The EL 10 NVIDIA repository provides no module stream — install the open kernel module driver directlysudo dnf install -y nvidia-openIf you prefer to install a specific driver version:
# List available versionssudo dnf list available nvidia-driver --showduplicates
# Install a specific version (replace <version> with one from the list above)sudo dnf install -y nvidia-driver-<version>Reboot after installation:
sudo rebootOption 2: ELRepo Repository
Section titled “Option 2: ELRepo Repository”ELRepo is a community-maintained driver repository for Enterprise Linux:
# Install the ELRepo repositorysudo dnf install -y https://www.elrepo.org/elrepo-release-9.el9.elrepo.noarch.rpm
# Install the NVIDIA driver (kmod version, auto-recompiles with kernel updates)sudo dnf install -y kmod-nvidia
# Or install the DKMS versionsudo dnf install -y nvidia-x11-drv# Install the ELRepo repositorysudo dnf install -y https://www.elrepo.org/elrepo-release-10.el10.elrepo.noarch.rpm
# Install the NVIDIA driver (kmod version)sudo dnf install -y kmod-nvidia
# Or install the DKMS versionsudo dnf install -y nvidia-x11-drvReboot:
sudo rebootVerify Driver Installation
Section titled “Verify Driver Installation”nvidia-smiExample successful output:
+-----------------------------------------------------------------------------------------+| NVIDIA-SMI 580.126.16 Driver Version: 580.126.16 CUDA Version: 13.3 ||-----------------------------------------+------------------------+----------------------+| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC || Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. ||=========================================+========================+======================|| 0 NVIDIA A100 80GB PCIe Off | 00000000:3B:00.0 Off | 0 || N/A 32C P0 43W / 300W | 0MiB / 81920MiB | 0% Default |+-----------------------------------------+------------------------+----------------------+Common nvidia-smi Commands
Section titled “Common nvidia-smi Commands”nvidia-smi is the NVIDIA System Management Interface tool for monitoring and managing GPUs.
# Basic informationnvidia-smi
# Continuous monitoring (refresh every second)watch -n 1 nvidia-smi
# View detailed GPU informationnvidia-smi -q
# View only GPU utilization and memorynvidia-smi --query-gpu=index,name,utilization.gpu,memory.used,memory.total --format=csv
# View processes running on the GPUnvidia-smi pmon -i 0
# Enable persistence mode (reduces first-call latency)sudo nvidia-smi -pm 1
# Set GPU power limitsudo nvidia-smi -pl 250
# View GPU temperaturenvidia-smi --query-gpu=temperature.gpu --format=csv,noheaderPersistence Mode Daemon
Section titled “Persistence Mode Daemon”On servers, it is recommended to enable nvidia-persistenced to avoid initialization delays on each GPU call:
sudo systemctl enable --now nvidia-persistencedInstall CUDA Toolkit
Section titled “Install CUDA Toolkit”The CUDA toolkit includes the compiler (nvcc), libraries, and development tools required for compiling GPU-accelerated programs.
Install via NVIDIA Repository
Section titled “Install via NVIDIA Repository”If you already added the NVIDIA CUDA repository when installing the driver, you can install directly:
# Install the latest CUDA toolkitsudo dnf install -y cuda-toolkit
# Or install a specific version (check which versions the repository provides first)sudo dnf list available 'cuda-toolkit-*'sudo dnf install -y cuda-toolkit-<version>Configure Environment Variables
Section titled “Configure Environment Variables”# Add CUDA to PATHcat >> ~/.bashrc <<'EOF'
# CUDAexport PATH=/usr/local/cuda/bin:$PATHexport LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATHEOF
source ~/.bashrcVerify CUDA Installation
Section titled “Verify CUDA Installation”nvcc --versionExample output:
nvcc: NVIDIA (R) Cuda compiler driverCopyright (c) 2005-2024 NVIDIA CorporationBuilt on ...Cuda compilation tools, release 13.3, V13.3.xxxCompile a Test Program
Section titled “Compile a Test Program”# Install CUDA samples (optional)sudo dnf install -y cuda-samples-12-6
# Compile the deviceQuery samplecd /usr/local/cuda/samples/1_Utilities/deviceQuery 2>/dev/null || \ echo "Samples may be in a different location; check /usr/local/cuda/extras/"
# Or write a test program manuallycat > /tmp/cuda_test.cu <<'EOF'#include <stdio.h>__global__ void hello() { printf("Hello from GPU thread %d\n", threadIdx.x);}int main() { hello<<<1, 5>>>(); cudaDeviceSynchronize(); return 0;}EOF
nvcc /tmp/cuda_test.cu -o /tmp/cuda_test && /tmp/cuda_testcuDNN Installation
Section titled “cuDNN Installation”Some deep learning frameworks require the cuDNN library:
# Install via NVIDIA repositorysudo dnf install -y cudnn9-cuda-12NVIDIA Container Toolkit
Section titled “NVIDIA Container Toolkit”The NVIDIA Container Toolkit enables Docker and Podman containers to access the host’s GPU, which is a key component for running AI applications in containers.
Install Container Toolkit
Section titled “Install Container Toolkit”# Add the NVIDIA Container Toolkit repositorycurl -s -L https://nvidia.github.io/libnvidia-container/stable/rpm/nvidia-container-toolkit.repo | \ sudo tee /etc/yum.repos.d/nvidia-container-toolkit.repo
# Installsudo dnf install -y nvidia-container-toolkitConfigure Docker
Section titled “Configure Docker”# Configure Docker runtimesudo nvidia-ctk runtime configure --runtime=docker
# Restart Dockersudo systemctl restart docker
# Verify GPU is available inside containerssudo docker run --rm --gpus all nvidia/cuda:13.3.1-base-ubi9 nvidia-smiConfigure Podman
Section titled “Configure Podman”# Configure Podman (CDI method)sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
# Verify CDI devicesnvidia-ctk cdi list
# Run a GPU containerpodman run --rm --device nvidia.com/gpu=all nvidia/cuda:13.3.1-base-ubi9 nvidia-smiUsing GPU in Docker Compose
Section titled “Using GPU in Docker Compose”Declare GPU resources in docker-compose.yml:
services: ai-app: image: your-ai-image:latest deploy: resources: reservations: devices: - driver: nvidia count: all # Or specify a number like 1 capabilities: [gpu]Driver Updates
Section titled “Driver Updates”Update Drivers
Section titled “Update Drivers”# Update via NVIDIA repositorysudo dnf update -y nvidia-driver cuda-toolkit
# Reboot after updatingsudo rebootRoll Back Drivers
Section titled “Roll Back Drivers”If the new driver has issues, you can roll back to a previous version:
# Check installed versionsrpm -qa | grep nvidia-driver
# Downgrade to a specific version (replace <version> with the target version)sudo dnf downgrade nvidia-driver-<version>sudo rebootnvidia-smi Error: NVIDIA-SMI has failed
Section titled “nvidia-smi Error: NVIDIA-SMI has failed”# Check if the driver module is loadedlsmod | grep nvidia
# If not loaded, try loading it manuallysudo modprobe nvidia
# Check kernel logssudo dmesg | grep -i nvidia
# The most common cause is Secure Boot blocking unsigned kernel modules# Disable Secure Boot in BIOS, or sign the modulesDriver Fails After Kernel Update
Section titled “Driver Fails After Kernel Update”If you used the DKMS installation method, the driver should auto-recompile after a kernel update. If it fails:
# Check DKMS statusdkms status
# Manually recompilesudo dkms autoinstall
# Rebootsudo rebootSpecifying GPUs in Multi-GPU Environments
Section titled “Specifying GPUs in Multi-GPU Environments”# List all GPUsnvidia-smi -L
# Specify visible GPUs via environment variableexport CUDA_VISIBLE_DEVICES=0 # Use only GPU 0export CUDA_VISIBLE_DEVICES=0,1 # Use GPU 0 and 1export CUDA_VISIBLE_DEVICES="" # Disable all GPUsCheck GPU-Compatible CUDA Version
Section titled “Check GPU-Compatible CUDA Version”# The top-right corner of nvidia-smi shows the maximum CUDA version supported by the drivernvidia-smi
# The actual installed CUDA versionnvcc --version