Skip to content

NVIDIA Drivers and CUDA

Running AI workloads on EL (Enterprise Linux) systems requires proper installation of NVIDIA GPU drivers and the CUDA toolkit. This guide covers driver installation, CUDA configuration, nvidia-smi usage, and NVIDIA Container Toolkit deployment.

Terminal window
lspci | grep -i nvidia

Example output:

3b:00.0 3D controller: NVIDIA Corporation A100 80GB PCIe (rev a1)

If there is no output, the system has not detected an NVIDIA GPU. Check the hardware connection or BIOS settings.

Terminal window
lspci -v -s $(lspci | grep -i nvidia | head -1 | awk '{print $1}')

There are two main ways to install drivers: through the NVIDIA official repository (recommended) or through ELRepo.

Ensure the system is updated and install necessary dependencies:

Terminal window
sudo dnf update -y
sudo dnf install -y kernel-devel kernel-headers gcc make dkms

Disable the nouveau open-source driver (if loaded):

Terminal window
# Check if nouveau is loaded
lsmod | grep nouveau
# If there is output, disable it
sudo tee /etc/modprobe.d/blacklist-nouveau.conf > /dev/null <<'EOF'
blacklist nouveau
options nouveau modeset=0
EOF
# Rebuild initramfs
sudo dracut --force
# Reboot
sudo reboot
Section titled “Option 1: NVIDIA Official Repository (Recommended)”

This is the recommended approach for getting the latest driver versions. Pick the repository that matches your EL release.

Add the NVIDIA CUDA repository and install the driver (EL 9)
# Add the NVIDIA CUDA repository (includes both drivers and CUDA)
sudo dnf config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel9/x86_64/cuda-rhel9.repo
# Clean cache
sudo dnf clean all
# Install the latest driver (DKMS module)
sudo dnf module install -y nvidia-driver:latest-dkms
Add the NVIDIA CUDA repository and install the driver (EL 10)
# Add the NVIDIA CUDA repository (rhel10)
sudo dnf config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel10/x86_64/cuda-rhel10.repo
# Clean cache
sudo dnf clean all
# The EL 10 NVIDIA repository provides no module stream — install the open kernel module driver directly
sudo dnf install -y nvidia-open

If you prefer to install a specific driver version:

List and install a specific driver version
# List available versions
sudo dnf list available nvidia-driver --showduplicates
# Install a specific version (replace <version> with one from the list above)
sudo dnf install -y nvidia-driver-<version>

Reboot after installation:

Terminal window
sudo reboot

ELRepo is a community-maintained driver repository for Enterprise Linux:

Install the NVIDIA driver via ELRepo (EL 9)
# Install the ELRepo repository
sudo dnf install -y https://www.elrepo.org/elrepo-release-9.el9.elrepo.noarch.rpm
# Install the NVIDIA driver (kmod version, auto-recompiles with kernel updates)
sudo dnf install -y kmod-nvidia
# Or install the DKMS version
sudo dnf install -y nvidia-x11-drv
Install the NVIDIA driver via ELRepo (EL 10)
# Install the ELRepo repository
sudo dnf install -y https://www.elrepo.org/elrepo-release-10.el10.elrepo.noarch.rpm
# Install the NVIDIA driver (kmod version)
sudo dnf install -y kmod-nvidia
# Or install the DKMS version
sudo dnf install -y nvidia-x11-drv

Reboot:

Terminal window
sudo reboot
Terminal window
nvidia-smi

Example successful output:

+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.126.16 Driver Version: 580.126.16 CUDA Version: 13.3 |
|-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
|=========================================+========================+======================|
| 0 NVIDIA A100 80GB PCIe Off | 00000000:3B:00.0 Off | 0 |
| N/A 32C P0 43W / 300W | 0MiB / 81920MiB | 0% Default |
+-----------------------------------------+------------------------+----------------------+

nvidia-smi is the NVIDIA System Management Interface tool for monitoring and managing GPUs.

Terminal window
# Basic information
nvidia-smi
# Continuous monitoring (refresh every second)
watch -n 1 nvidia-smi
# View detailed GPU information
nvidia-smi -q
# View only GPU utilization and memory
nvidia-smi --query-gpu=index,name,utilization.gpu,memory.used,memory.total --format=csv
# View processes running on the GPU
nvidia-smi pmon -i 0
# Enable persistence mode (reduces first-call latency)
sudo nvidia-smi -pm 1
# Set GPU power limit
sudo nvidia-smi -pl 250
# View GPU temperature
nvidia-smi --query-gpu=temperature.gpu --format=csv,noheader

On servers, it is recommended to enable nvidia-persistenced to avoid initialization delays on each GPU call:

Terminal window
sudo systemctl enable --now nvidia-persistenced

The CUDA toolkit includes the compiler (nvcc), libraries, and development tools required for compiling GPU-accelerated programs.

If you already added the NVIDIA CUDA repository when installing the driver, you can install directly:

Install the CUDA toolkit
# Install the latest CUDA toolkit
sudo dnf install -y cuda-toolkit
# Or install a specific version (check which versions the repository provides first)
sudo dnf list available 'cuda-toolkit-*'
sudo dnf install -y cuda-toolkit-<version>
Terminal window
# Add CUDA to PATH
cat >> ~/.bashrc <<'EOF'
# CUDA
export PATH=/usr/local/cuda/bin:$PATH
export LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATH
EOF
source ~/.bashrc
Terminal window
nvcc --version

Example output:

nvcc: NVIDIA (R) Cuda compiler driver
Copyright (c) 2005-2024 NVIDIA Corporation
Built on ...
Cuda compilation tools, release 13.3, V13.3.xxx
Terminal window
# Install CUDA samples (optional)
sudo dnf install -y cuda-samples-12-6
# Compile the deviceQuery sample
cd /usr/local/cuda/samples/1_Utilities/deviceQuery 2>/dev/null || \
echo "Samples may be in a different location; check /usr/local/cuda/extras/"
# Or write a test program manually
cat > /tmp/cuda_test.cu <<'EOF'
#include <stdio.h>
__global__ void hello() {
printf("Hello from GPU thread %d\n", threadIdx.x);
}
int main() {
hello<<<1, 5>>>();
cudaDeviceSynchronize();
return 0;
}
EOF
nvcc /tmp/cuda_test.cu -o /tmp/cuda_test && /tmp/cuda_test

Some deep learning frameworks require the cuDNN library:

Terminal window
# Install via NVIDIA repository
sudo dnf install -y cudnn9-cuda-12

The NVIDIA Container Toolkit enables Docker and Podman containers to access the host’s GPU, which is a key component for running AI applications in containers.

Terminal window
# Add the NVIDIA Container Toolkit repository
curl -s -L https://nvidia.github.io/libnvidia-container/stable/rpm/nvidia-container-toolkit.repo | \
sudo tee /etc/yum.repos.d/nvidia-container-toolkit.repo
# Install
sudo dnf install -y nvidia-container-toolkit
Terminal window
# Configure Docker runtime
sudo nvidia-ctk runtime configure --runtime=docker
# Restart Docker
sudo systemctl restart docker
# Verify GPU is available inside containers
sudo docker run --rm --gpus all nvidia/cuda:13.3.1-base-ubi9 nvidia-smi
Terminal window
# Configure Podman (CDI method)
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
# Verify CDI devices
nvidia-ctk cdi list
# Run a GPU container
podman run --rm --device nvidia.com/gpu=all nvidia/cuda:13.3.1-base-ubi9 nvidia-smi

Declare GPU resources in docker-compose.yml:

services:
ai-app:
image: your-ai-image:latest
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all # Or specify a number like 1
capabilities: [gpu]
Terminal window
# Update via NVIDIA repository
sudo dnf update -y nvidia-driver cuda-toolkit
# Reboot after updating
sudo reboot

If the new driver has issues, you can roll back to a previous version:

Roll back to a specific version
# Check installed versions
rpm -qa | grep nvidia-driver
# Downgrade to a specific version (replace <version> with the target version)
sudo dnf downgrade nvidia-driver-<version>
sudo reboot
Terminal window
# Check if the driver module is loaded
lsmod | grep nvidia
# If not loaded, try loading it manually
sudo modprobe nvidia
# Check kernel logs
sudo dmesg | grep -i nvidia
# The most common cause is Secure Boot blocking unsigned kernel modules
# Disable Secure Boot in BIOS, or sign the modules

If you used the DKMS installation method, the driver should auto-recompile after a kernel update. If it fails:

Terminal window
# Check DKMS status
dkms status
# Manually recompile
sudo dkms autoinstall
# Reboot
sudo reboot
Terminal window
# List all GPUs
nvidia-smi -L
# Specify visible GPUs via environment variable
export CUDA_VISIBLE_DEVICES=0 # Use only GPU 0
export CUDA_VISIBLE_DEVICES=0,1 # Use GPU 0 and 1
export CUDA_VISIBLE_DEVICES="" # Disable all GPUs
Terminal window
# The top-right corner of nvidia-smi shows the maximum CUDA version supported by the driver
nvidia-smi
# The actual installed CUDA version
nvcc --version