Skip to content

llama.cpp

llama.cpp is a large language model inference framework written in pure C/C++, created by Georgi Gerganov. Its greatest advantage is that it runs without a Python environment, supports both CPU and GPU inference, and can efficiently run quantized models in GGUF format. For resource-constrained servers or scenarios requiring maximum performance, llama.cpp is an excellent choice.

  • CentOS Stream 9 / AlmaLinux 9 / Rocky Linux 9
  • GCC 11+ or Clang 14+
  • CMake 3.21+
  • At least 4 GB RAM (for running 7B quantized models)
  • Optional: NVIDIA GPU + CUDA (for GPU acceleration)
Terminal window
sudo dnf install -y git gcc gcc-c++ make cmake

If CUDA support is needed, make sure the NVIDIA driver and CUDA toolkit are installed (see NVIDIA Drivers and CUDA).

Terminal window
cd /opt
sudo mkdir -p llama-cpp && sudo chown $(whoami):$(whoami) llama-cpp
git clone https://github.com/ggerganov/llama.cpp.git /opt/llama-cpp
cd /opt/llama-cpp
Terminal window
cd /opt/llama-cpp
cmake -B build
cmake --build build --config Release -j$(nproc)

After compilation, the binaries are located in the build/bin/ directory.

Terminal window
cd /opt/llama-cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j$(nproc)

Verify the CUDA build succeeded:

Terminal window
./build/bin/llama-cli --version
Terminal window
# Enable OpenBLAS acceleration (CPU matrix operation optimization)
sudo dnf install -y openblas-devel
cmake -B build -DGGML_BLAS=ON -DGGML_BLAS_VENDOR=OpenBLAS
cmake --build build --config Release -j$(nproc)
# Enable Vulkan (for AMD GPUs or NVIDIA GPUs without CUDA)
sudo dnf install -y vulkan-headers vulkan-loader-devel
cmake -B build -DGGML_VULKAN=ON
cmake --build build --config Release -j$(nproc)

GGUF (GPT-Generated Unified Format) is the model format used by llama.cpp. It supports various quantization levels, striking a balance between model size and inference quality.

Download pre-quantized GGUF models from Hugging Face:

Terminal window
mkdir -p /opt/llama-cpp/models
# Download Qwen2.5-7B-Instruct GGUF (Q4_K_M quantization, approximately 4.7 GB)
cd /opt/llama-cpp/models
wget https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-GGUF/resolve/main/qwen2.5-7b-instruct-q4_k_m.gguf
# Download Llama 3.1 8B GGUF
wget https://huggingface.co/bartowski/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf
Quantization TypeSize (7B Model)QualitySpeed
Q2_K~2.7 GBLowerFastest
Q3_K_M~3.3 GBFairFast
Q4_K_M~4.7 GBGood (Recommended)Moderately fast
Q5_K_M~5.3 GBVery goodMedium
Q6_K~5.9 GBExcellentModerately slow
Q8_0~7.2 GBNear originalSlow
F16~14 GBOriginal precisionSlowest

Q4_K_M is generally recommended as it provides a good balance between quality and size.

Terminal window
cd /opt/llama-cpp
./build/bin/llama-cli \
-m models/qwen2.5-7b-instruct-q4_k_m.gguf \
-c 4096 \
-n 512 \
--chat-template chatml \
-cnv

Parameter descriptions:

ParameterDescription
-mModel file path
-cContext length
-nMaximum number of tokens to generate
--chat-templateChat template format
-cnvEnable conversation mode
-tNumber of threads (auto-detected by default)
-nglNumber of layers to offload to GPU (available with CUDA builds)

If compiled with CUDA, you can offload some or all model layers to the GPU:

Terminal window
# Offload all layers to GPU (requires sufficient GPU VRAM)
./build/bin/llama-cli \
-m models/qwen2.5-7b-instruct-q4_k_m.gguf \
-c 4096 \
-ngl 99 \
-cnv
# Offload only some layers (when VRAM is limited)
./build/bin/llama-cli \
-m models/qwen2.5-7b-instruct-q4_k_m.gguf \
-c 4096 \
-ngl 20 \
-cnv
Terminal window
./build/bin/llama-cli \
-m models/qwen2.5-7b-instruct-q4_k_m.gguf \
-p "Explain what a Linux kernel is in simple terms" \
-n 256 \
-c 2048
Terminal window
./build/bin/llama-bench \
-m models/qwen2.5-7b-instruct-q4_k_m.gguf \
-t $(nproc)

llama.cpp includes a built-in HTTP server that provides an OpenAI API-compatible interface.

Terminal window
cd /opt/llama-cpp
./build/bin/llama-server \
-m models/qwen2.5-7b-instruct-q4_k_m.gguf \
--host 0.0.0.0 \
--port 8080 \
-c 4096 \
-ngl 99 \
--chat-template chatml

The server provides an OpenAI-compatible API:

Terminal window
# Chat Completions
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen2.5",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is SELinux?"}
],
"temperature": 0.7,
"max_tokens": 512
}'
Terminal window
# Text Completions
curl http://localhost:8080/v1/completions \
-H "Content-Type: application/json" \
-d '{
"prompt": "The advantages of the Linux operating system include",
"max_tokens": 256
}'
Terminal window
# Health check
curl http://localhost:8080/health

The server also includes a simple web interface accessible by visiting http://your-server-ip:8080 in your browser.

Terminal window
sudo tee /etc/systemd/system/llama-server.service > /dev/null <<'EOF'
[Unit]
Description=llama.cpp Server
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=root
WorkingDirectory=/opt/llama-cpp
ExecStart=/opt/llama-cpp/build/bin/llama-server \
-m /opt/llama-cpp/models/qwen2.5-7b-instruct-q4_k_m.gguf \
--host 0.0.0.0 \
--port 8080 \
-c 4096 \
-ngl 99 \
--chat-template chatml
Restart=on-failure
RestartSec=10
LimitNOFILE=65536
[Install]
WantedBy=multi-user.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now llama-server

If you have a raw Hugging Face model, you can convert and quantize it yourself using llama.cpp.

Terminal window
cd /opt/llama-cpp
# Install Python dependencies (only needed for conversion)
pip install -r requirements.txt
# Convert a Hugging Face model to GGUF (F16 precision)
python convert_hf_to_gguf.py /path/to/huggingface-model/ \
--outfile models/my-model-f16.gguf \
--outtype f16
Terminal window
# Quantize an F16 model to Q4_K_M
./build/bin/llama-quantize \
models/my-model-f16.gguf \
models/my-model-q4_k_m.gguf \
Q4_K_M
# Quantize to Q5_K_M (higher quality)
./build/bin/llama-quantize \
models/my-model-f16.gguf \
models/my-model-q5_k_m.gguf \
Q5_K_M
Terminal window
# Quick test to verify the quantized model works
./build/bin/llama-cli \
-m models/my-model-q4_k_m.gguf \
-p "Hello, how are you?" \
-n 64
Terminal window
sudo firewall-cmd --permanent --add-port=8080/tcp
sudo firewall-cmd --reload

llama.cpp is updated very frequently. It is recommended to update regularly for performance improvements and new model support:

Terminal window
cd /opt/llama-cpp
git pull
# Clean old build files and recompile
rm -rf build
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j$(nproc)
# Restart the service
sudo systemctl restart llama-server
Terminal window
# Install a newer version of CMake
sudo dnf install -y cmake3
# Or compile CMake from source
Terminal window
# Check if all CPU cores are being used
./build/bin/llama-cli -m model.gguf -t $(nproc) ...
# If you have a GPU, make sure you compiled with CUDA and set -ngl
./build/bin/llama-cli -m model.gguf -ngl 99 ...
# Use a smaller quantized model
# Q4_K_M is typically 40-50% faster than Q8_0
Terminal window
# Use a smaller quantization level
# Q2_K is about 1/5 the size of F16
# Reduce context length
./build/bin/llama-cli -m model.gguf -c 2048 ...
# Use mmap (enabled by default) to avoid loading the entire model into memory
Terminal window
# Verify CUDA toolkit version
nvcc --version
# Verify CUDA was enabled during compilation
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j$(nproc)
# Check if the GPU is detected
nvidia-smi