AI on EL Deployment Overview
Deploying AI inference services on EL systems requires a complete stack from hardware drivers to the application layer. This page helps you understand the architecture and choose the right tool combination.
Deployment Stack
Section titled “Deployment Stack”Hardware GPU (NVIDIA) ↓Drivers NVIDIA Driver + CUDA Toolkit ↓Runtime Ollama / vLLM / llama.cpp ↓Models LLaMA / Qwen / Gemma / Mistral (GGUF/HF format) ↓API Layer OpenAI-compatible API (built-in) ↓Frontend Open WebUI / Custom apps ↓Gateway Nginx reverse proxy + SSLChoosing a Solution
Section titled “Choosing a Solution”| Solution | Best For | Difficulty | GPU Required |
|---|---|---|---|
| Ollama | Personal use, development | Easy | Optional |
| vLLM | Production inference, high concurrency | Medium | Required |
| llama.cpp | Maximum optimization, CPU inference, quantization | Medium | Optional |
| Open WebUI + Ollama | Team ChatGPT alternative | Easy | Optional |
| Stable Diffusion | Image generation | Medium | Required |
Recommended Deployment Order
Section titled “Recommended Deployment Order”Quick Start (30 minutes)
Section titled “Quick Start (30 minutes)”Just want to run a local LLM quickly:
- Install Ollama — one command install
ollama run llama3.1— start chatting
Standard Deployment (2-4 hours)
Section titled “Standard Deployment (2-4 hours)”Provide a usable AI service for your team:
- NVIDIA Drivers & CUDA — install GPU drivers
- Ollama — install and download models
- Open WebUI — deploy web chat interface
- Nginx Reverse Proxy — configure HTTPS and domain
- Firewall — configure access control
Production Deployment
Section titled “Production Deployment”High-concurrency, low-latency inference service:
- NVIDIA Drivers & CUDA — GPU drivers and CUDA
- vLLM — high-performance inference engine
- Nginx Reverse Proxy — load balancing and SSL
- Monitoring — GPU and service monitoring
- Performance Tuning — kernel and network optimization
EL-Specific Considerations
Section titled “EL-Specific Considerations”- SELinux: For containerized deployments, set
setsebool -P container_use_devices onto allow GPU access - Firewall: AI services listen on high ports (8080/11434 etc.) — open in firewalld or use Nginx reverse proxy
- systemd: Register Ollama/vLLM as systemd services for auto-start and crash recovery
- Storage: Large model files are typically 5-70 GB — ensure sufficient space in
/varor model directory - NVIDIA driver compatibility: On EL 9, use the NVIDIA official repo to avoid kernel version conflicts
Common issue: GPU not available
Section titled “Common issue: GPU not available”The most common post-deployment problem is “nvidia-smi works, but Ollama/vLLM can’t see the GPU and falls back to CPU.” Troubleshoot in this order:
-
Confirm the host sees the GPU
Terminal window nvidia-smi # if no GPU is listed, the driver isn't installed — go back to the NVIDIA driver doc -
Containerized deployment: check the NVIDIA Container Toolkit
Skip this for bare-metal installs. If the GPU is invisible inside containers,
nvidia-container-toolkitis usually missing or unconfigured:Terminal window # Generate CDI device config after install (CDI is recommended for Podman)sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml# Mount the GPU at runtimepodman run --rm --device nvidia.com/gpu=all nvidia/cuda:12.4.0-base-ubi9 nvidia-smi -
SELinux blocking device access
Terminal window # Allow containers to access GPU device nodes under /devsudo setsebool -P container_use_devices on# When troubleshooting, inspect AVC denialssudo ausearch -m avc -ts recent -
CUDA / driver version mismatch
If vLLM/PyTorch reports
CUDA driver version is insufficient, the driver is older than the CUDA runtime the framework needs. The CUDA version shown in the top-right ofnvidia-smiis the highest the driver supports and must be ≥ what the framework requires. -
Out of memory (OOM)
If vLLM fails with
CUDA out of memory: lower--gpu-memory-utilization(default 0.9), reduce--max-model-len, or for large models use--tensor-parallel-size Nto shard across multiple GPUs.
Further Reading
Section titled “Further Reading”- NVIDIA Drivers & CUDA — GPU driver installation
- Ollama — simplest way to get started
- Claude Code — AI coding assistant