vLLM Inference Service
vLLM is a high-performance large language model inference and serving framework developed at UC Berkeley. Through innovations such as PagedAttention, it significantly improves inference throughput, achieving several times the performance of traditional approaches. vLLM natively provides an OpenAI API-compatible HTTP service, making it suitable for deploying high-concurrency AI inference services in production environments.
System Requirements
Section titled “System Requirements”- CentOS Stream 9 & 10 / AlmaLinux 9.x & 10.x / Rocky Linux 9.x & 10.x
- Python 3.10 - 3.14 (EL 10 defaults to Python 3.12; see the
requires-pythonof your vLLM version) - NVIDIA GPU (compute capability 7.0 or above, i.e., V100/T4/A100/RTX 20 series and newer)
- NVIDIA driver: R570 or newer + CUDA 12.8 or newer is recommended (the current default wheels are built for CUDA 12.9; CUDA 13 requires an R580+ driver, and Blackwell GPUs need at least CUDA 12.8; see NVIDIA Drivers and CUDA)
- At least 16 GB GPU VRAM (for running 7B models)
Pre-Installation
Section titled “Pre-Installation”Install System Dependencies
Section titled “Install System Dependencies”sudo dnf install -y python3.12 python3.12-devel python3.12-pip gcc gcc-c++ makeIf Python 3.12 is not available by default (mainly on EL 9), install it from the CRB repository:
sudo dnf install -y epel-releasesudo dnf config-manager --set-enabled crbsudo dnf install -y python3.12 python3.12-devel python3.12-pipVerify GPU Environment
Section titled “Verify GPU Environment”# Confirm the NVIDIA driver is workingnvidia-smi
# Confirm CUDA versionnvcc --versionCreate a Python Virtual Environment
Section titled “Create a Python Virtual Environment”It is strongly recommended to install vLLM in a virtual environment to avoid conflicts with system Python packages:
# Create project directorysudo mkdir -p /opt/vllmsudo chown $(whoami):$(whoami) /opt/vllm
# Create virtual environmentpython3.12 -m venv /opt/vllm/venv
# Activate virtual environmentsource /opt/vllm/venv/bin/activate
# Upgrade pippip install --upgrade pip setuptools wheelInstall vLLM
Section titled “Install vLLM”Install via pip (Recommended)
Section titled “Install via pip (Recommended)”source /opt/vllm/venv/bin/activatepip install vllmThe installation will automatically pull dependencies such as PyTorch and Triton, which may take several minutes.
Verify Installation
Section titled “Verify Installation”source /opt/vllm/venv/bin/activatepython -c "import vllm; print(vllm.__version__)"Download Models
Section titled “Download Models”vLLM uses Hugging Face format models. You can download models in advance or let vLLM download them automatically at startup.
source /opt/vllm/venv/bin/activatepip install huggingface_hub[cli]
# Download Qwen2.5-7B-Instruct to a specified directoryhuggingface-cli download Qwen/Qwen2.5-7B-Instruct \ --local-dir /opt/vllm/models/Qwen2.5-7B-Instruct
# Or download Meta-Llama-3.1-8B-Instruct (gated model: request access on HF and log in first)huggingface-cli download meta-llama/Meta-Llama-3.1-8B-Instruct \ --local-dir /opt/vllm/models/Meta-Llama-3.1-8B-InstructIf Hugging Face access is slow, you can set a mirror:
export HF_ENDPOINT=https://hf-mirror.comhuggingface-cli download Qwen/Qwen2.5-7B-Instruct \ --local-dir /opt/vllm/models/Qwen2.5-7B-InstructRun OpenAI-Compatible API Service
Section titled “Run OpenAI-Compatible API Service”vLLM includes a built-in OpenAI API-compatible HTTP server:
source /opt/vllm/venv/bin/activate
vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct \ --host 0.0.0.0 \ --port 8000 \ --served-model-name qwen2.5-7bOnce the service is running, you can call it using the standard OpenAI API format:
# Test Chat Completions endpointcurl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "qwen2.5-7b", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is the difference between CentOS Stream and AlmaLinux?"} ], "temperature": 0.7, "max_tokens": 512 }'# List available modelscurl http://localhost:8000/v1/models# Test Completions endpointcurl http://localhost:8000/v1/completions \ -H "Content-Type: application/json" \ -d '{ "model": "qwen2.5-7b", "prompt": "The advantages of Linux are", "max_tokens": 128 }'Performance Tuning
Section titled “Performance Tuning”Key Startup Parameters
Section titled “Key Startup Parameters”vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct \ --host 0.0.0.0 \ --port 8000 \ --served-model-name qwen2.5-7b \ --tensor-parallel-size 2 \ --gpu-memory-utilization 0.90 \ --max-model-len 8192 \ --max-num-seqs 64 \ --enable-prefix-caching \ --dtype autoParameter descriptions:
| Parameter | Description |
|---|---|
--tensor-parallel-size | Tensor parallelism degree; set to the number of GPUs. Used with multiple GPUs |
--gpu-memory-utilization | GPU VRAM utilization ratio; default is 0.9, should not exceed 0.95 |
--max-model-len | Maximum context length; reducing this value lowers VRAM usage |
--max-num-seqs | Maximum number of concurrent sequences |
--enable-prefix-caching | Enable prefix caching; significantly speeds up similar requests |
--dtype | Data type; auto selects automatically, or specify float16 or bfloat16 |
--quantization | Quantization method; supports awq, gptq, squeezellm, etc. |
Use Quantized Models to Reduce VRAM Usage
Section titled “Use Quantized Models to Reduce VRAM Usage”# Use an AWQ quantized modelvllm serve Qwen/Qwen2.5-7B-Instruct-AWQ \ --quantization awq \ --host 0.0.0.0 \ --port 8000Multi-GPU Inference
Section titled “Multi-GPU Inference”# Use 2 GPUs for tensor parallelismvllm serve /opt/vllm/models/Qwen2.5-7B-Instruct \ --tensor-parallel-size 2 \ --host 0.0.0.0 \ --port 8000systemd Service
Section titled “systemd Service”Create a systemd service file for auto-start and process management:
sudo tee /etc/systemd/system/vllm.service > /dev/null <<'EOF'[Unit]Description=vLLM OpenAI Compatible API ServerAfter=network-online.targetWants=network-online.target
[Service]Type=simpleUser=rootGroup=rootWorkingDirectory=/opt/vllmExecStart=/opt/vllm/venv/bin/vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct \ --host 0.0.0.0 \ --port 8000 \ --served-model-name qwen2.5-7b \ --gpu-memory-utilization 0.90 \ --max-model-len 8192Restart=on-failureRestartSec=10Environment="CUDA_VISIBLE_DEVICES=0"Environment="HF_HOME=/opt/vllm/hf_cache"
# Resource limitsLimitNOFILE=65536LimitNPROC=65536
[Install]WantedBy=multi-user.targetEOFEnable and start the service:
sudo systemctl daemon-reloadsudo systemctl enable --now vllmManage the service:
# Check statussudo systemctl status vllm
# View logssudo journalctl -u vllm -f
# Restartsudo systemctl restart vllmFirewall Configuration
Section titled “Firewall Configuration”sudo firewall-cmd --permanent --add-port=8000/tcpsudo firewall-cmd --reloadUsing the Python Client
Section titled “Using the Python Client”Since vLLM is compatible with the OpenAI API, you can use the OpenAI Python client directly:
pip install openaifrom openai import OpenAI
client = OpenAI( base_url="http://localhost:8000/v1", api_key="not-needed", # vLLM does not require an API key by default)
response = client.chat.completions.create( model="qwen2.5-7b", messages=[ {"role": "system", "content": "You are a Linux expert."}, {"role": "user", "content": "How do I configure the firewall on AlmaLinux 9?"}, ], temperature=0.7, max_tokens=1024,)
print(response.choices[0].message.content)CUDA Out of Memory
Section titled “CUDA Out of Memory”If you encounter a CUDA out of memory error:
# Reduce the maximum context lengthvllm serve /opt/vllm/models/Qwen2.5-7B-Instruct --max-model-len 4096
# Or lower the GPU VRAM utilizationvllm serve /opt/vllm/models/Qwen2.5-7B-Instruct --gpu-memory-utilization 0.80
# Or use a quantized modelvllm serve Qwen/Qwen2.5-7B-Instruct-AWQ --quantization awqModel Download Timeout
Section titled “Model Download Timeout”Set the Hugging Face mirror environment variable:
sudo systemctl edit vllm[Service]Environment="HF_ENDPOINT=https://hf-mirror.com"Monitor GPU Usage
Section titled “Monitor GPU Usage”Monitor the GPU while the service is running:
watch -n 1 nvidia-smi