Skip to content

vLLM Inference Service

vLLM is a high-performance large language model inference and serving framework developed at UC Berkeley. Through innovations such as PagedAttention, it significantly improves inference throughput, achieving several times the performance of traditional approaches. vLLM natively provides an OpenAI API-compatible HTTP service, making it suitable for deploying high-concurrency AI inference services in production environments.

  • CentOS Stream 9 & 10 / AlmaLinux 9.x & 10.x / Rocky Linux 9.x & 10.x
  • Python 3.10 - 3.14 (EL 10 defaults to Python 3.12; see the requires-python of your vLLM version)
  • NVIDIA GPU (compute capability 7.0 or above, i.e., V100/T4/A100/RTX 20 series and newer)
  • NVIDIA driver: R570 or newer + CUDA 12.8 or newer is recommended (the current default wheels are built for CUDA 12.9; CUDA 13 requires an R580+ driver, and Blackwell GPUs need at least CUDA 12.8; see NVIDIA Drivers and CUDA)
  • At least 16 GB GPU VRAM (for running 7B models)
Terminal window
sudo dnf install -y python3.12 python3.12-devel python3.12-pip gcc gcc-c++ make

If Python 3.12 is not available by default (mainly on EL 9), install it from the CRB repository:

EL 9: enable CRB and install Python 3.12
sudo dnf install -y epel-release
sudo dnf config-manager --set-enabled crb
sudo dnf install -y python3.12 python3.12-devel python3.12-pip
Terminal window
# Confirm the NVIDIA driver is working
nvidia-smi
# Confirm CUDA version
nvcc --version

It is strongly recommended to install vLLM in a virtual environment to avoid conflicts with system Python packages:

Terminal window
# Create project directory
sudo mkdir -p /opt/vllm
sudo chown $(whoami):$(whoami) /opt/vllm
# Create virtual environment
python3.12 -m venv /opt/vllm/venv
# Activate virtual environment
source /opt/vllm/venv/bin/activate
# Upgrade pip
pip install --upgrade pip setuptools wheel
Terminal window
source /opt/vllm/venv/bin/activate
pip install vllm

The installation will automatically pull dependencies such as PyTorch and Triton, which may take several minutes.

Terminal window
source /opt/vllm/venv/bin/activate
python -c "import vllm; print(vllm.__version__)"

vLLM uses Hugging Face format models. You can download models in advance or let vLLM download them automatically at startup.

Terminal window
source /opt/vllm/venv/bin/activate
pip install huggingface_hub[cli]
# Download Qwen2.5-7B-Instruct to a specified directory
huggingface-cli download Qwen/Qwen2.5-7B-Instruct \
--local-dir /opt/vllm/models/Qwen2.5-7B-Instruct
# Or download Meta-Llama-3.1-8B-Instruct (gated model: request access on HF and log in first)
huggingface-cli download meta-llama/Meta-Llama-3.1-8B-Instruct \
--local-dir /opt/vllm/models/Meta-Llama-3.1-8B-Instruct

If Hugging Face access is slow, you can set a mirror:

Terminal window
export HF_ENDPOINT=https://hf-mirror.com
huggingface-cli download Qwen/Qwen2.5-7B-Instruct \
--local-dir /opt/vllm/models/Qwen2.5-7B-Instruct

vLLM includes a built-in OpenAI API-compatible HTTP server:

Terminal window
source /opt/vllm/venv/bin/activate
vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 8000 \
--served-model-name qwen2.5-7b

Once the service is running, you can call it using the standard OpenAI API format:

Terminal window
# Test Chat Completions endpoint
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen2.5-7b",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the difference between CentOS Stream and AlmaLinux?"}
],
"temperature": 0.7,
"max_tokens": 512
}'
Terminal window
# List available models
curl http://localhost:8000/v1/models
Terminal window
# Test Completions endpoint
curl http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen2.5-7b",
"prompt": "The advantages of Linux are",
"max_tokens": 128
}'
Terminal window
vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 8000 \
--served-model-name qwen2.5-7b \
--tensor-parallel-size 2 \
--gpu-memory-utilization 0.90 \
--max-model-len 8192 \
--max-num-seqs 64 \
--enable-prefix-caching \
--dtype auto

Parameter descriptions:

ParameterDescription
--tensor-parallel-sizeTensor parallelism degree; set to the number of GPUs. Used with multiple GPUs
--gpu-memory-utilizationGPU VRAM utilization ratio; default is 0.9, should not exceed 0.95
--max-model-lenMaximum context length; reducing this value lowers VRAM usage
--max-num-seqsMaximum number of concurrent sequences
--enable-prefix-cachingEnable prefix caching; significantly speeds up similar requests
--dtypeData type; auto selects automatically, or specify float16 or bfloat16
--quantizationQuantization method; supports awq, gptq, squeezellm, etc.
Terminal window
# Use an AWQ quantized model
vllm serve Qwen/Qwen2.5-7B-Instruct-AWQ \
--quantization awq \
--host 0.0.0.0 \
--port 8000
Terminal window
# Use 2 GPUs for tensor parallelism
vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct \
--tensor-parallel-size 2 \
--host 0.0.0.0 \
--port 8000

Create a systemd service file for auto-start and process management:

Terminal window
sudo tee /etc/systemd/system/vllm.service > /dev/null <<'EOF'
[Unit]
Description=vLLM OpenAI Compatible API Server
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=root
Group=root
WorkingDirectory=/opt/vllm
ExecStart=/opt/vllm/venv/bin/vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 8000 \
--served-model-name qwen2.5-7b \
--gpu-memory-utilization 0.90 \
--max-model-len 8192
Restart=on-failure
RestartSec=10
Environment="CUDA_VISIBLE_DEVICES=0"
Environment="HF_HOME=/opt/vllm/hf_cache"
# Resource limits
LimitNOFILE=65536
LimitNPROC=65536
[Install]
WantedBy=multi-user.target
EOF

Enable and start the service:

Terminal window
sudo systemctl daemon-reload
sudo systemctl enable --now vllm

Manage the service:

Terminal window
# Check status
sudo systemctl status vllm
# View logs
sudo journalctl -u vllm -f
# Restart
sudo systemctl restart vllm
Terminal window
sudo firewall-cmd --permanent --add-port=8000/tcp
sudo firewall-cmd --reload

Since vLLM is compatible with the OpenAI API, you can use the OpenAI Python client directly:

Terminal window
pip install openai
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed", # vLLM does not require an API key by default
)
response = client.chat.completions.create(
model="qwen2.5-7b",
messages=[
{"role": "system", "content": "You are a Linux expert."},
{"role": "user", "content": "How do I configure the firewall on AlmaLinux 9?"},
],
temperature=0.7,
max_tokens=1024,
)
print(response.choices[0].message.content)

If you encounter a CUDA out of memory error:

Terminal window
# Reduce the maximum context length
vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct --max-model-len 4096
# Or lower the GPU VRAM utilization
vllm serve /opt/vllm/models/Qwen2.5-7B-Instruct --gpu-memory-utilization 0.80
# Or use a quantized model
vllm serve Qwen/Qwen2.5-7B-Instruct-AWQ --quantization awq

Set the Hugging Face mirror environment variable:

Terminal window
sudo systemctl edit vllm
[Service]
Environment="HF_ENDPOINT=https://hf-mirror.com"

Monitor the GPU while the service is running:

Terminal window
watch -n 1 nvidia-smi