Skip to content

AI on EL Deployment Overview

Deploying AI inference services on EL systems requires a complete stack from hardware drivers to the application layer. This page helps you understand the architecture and choose the right tool combination.

Hardware GPU (NVIDIA)
↓
Drivers NVIDIA Driver + CUDA Toolkit
↓
Runtime Ollama / vLLM / llama.cpp
↓
Models LLaMA / Qwen / Gemma / Mistral (GGUF/HF format)
↓
API Layer OpenAI-compatible API (built-in)
↓
Frontend Open WebUI / Custom apps
↓
Gateway Nginx reverse proxy + SSL
SolutionBest ForDifficultyGPU Required
OllamaPersonal use, developmentEasyOptional
vLLMProduction inference, high concurrencyMediumRequired
llama.cppMaximum optimization, CPU inference, quantizationMediumOptional
Open WebUI + OllamaTeam ChatGPT alternativeEasyOptional
Stable DiffusionImage generationMediumRequired

Just want to run a local LLM quickly:

  1. Install Ollama — one command install
  2. ollama run llama3.1 — start chatting

Provide a usable AI service for your team:

  1. NVIDIA Drivers & CUDA — install GPU drivers
  2. Ollama — install and download models
  3. Open WebUI — deploy web chat interface
  4. Nginx Reverse Proxy — configure HTTPS and domain
  5. Firewall — configure access control

High-concurrency, low-latency inference service:

  1. NVIDIA Drivers & CUDA — GPU drivers and CUDA
  2. vLLM — high-performance inference engine
  3. Nginx Reverse Proxy — load balancing and SSL
  4. Monitoring — GPU and service monitoring
  5. Performance Tuning — kernel and network optimization
  • SELinux: For containerized deployments, set setsebool -P container_use_devices on to allow GPU access
  • Firewall: AI services listen on high ports (8080/11434 etc.) — open in firewalld or use Nginx reverse proxy
  • systemd: Register Ollama/vLLM as systemd services for auto-start and crash recovery
  • Storage: Large model files are typically 5-70 GB — ensure sufficient space in /var or model directory
  • NVIDIA driver compatibility: On EL 9, use the NVIDIA official repo to avoid kernel version conflicts

The most common post-deployment problem is “nvidia-smi works, but Ollama/vLLM can’t see the GPU and falls back to CPU.” Troubleshoot in this order:

  1. Confirm the host sees the GPU

    Terminal window
    nvidia-smi # if no GPU is listed, the driver isn't installed — go back to the NVIDIA driver doc
  2. Containerized deployment: check the NVIDIA Container Toolkit

    Skip this for bare-metal installs. If the GPU is invisible inside containers, nvidia-container-toolkit is usually missing or unconfigured:

    Terminal window
    # Generate CDI device config after install (CDI is recommended for Podman)
    sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
    # Mount the GPU at runtime
    podman run --rm --device nvidia.com/gpu=all nvidia/cuda:12.4.0-base-ubi9 nvidia-smi
  3. SELinux blocking device access

    Terminal window
    # Allow containers to access GPU device nodes under /dev
    sudo setsebool -P container_use_devices on
    # When troubleshooting, inspect AVC denials
    sudo ausearch -m avc -ts recent
  4. CUDA / driver version mismatch

    If vLLM/PyTorch reports CUDA driver version is insufficient, the driver is older than the CUDA runtime the framework needs. The CUDA version shown in the top-right of nvidia-smi is the highest the driver supports and must be ≥ what the framework requires.

  5. Out of memory (OOM)

    If vLLM fails with CUDA out of memory: lower --gpu-memory-utilization (default 0.9), reduce --max-model-len, or for large models use --tensor-parallel-size N to shard across multiple GPUs.