llama.cpp
llama.cpp is a large language model inference framework written in pure C/C++, created by Georgi Gerganov. Its greatest advantage is that it runs without a Python environment, supports both CPU and GPU inference, and can efficiently run quantized models in GGUF format. For resource-constrained servers or scenarios requiring maximum performance, llama.cpp is an excellent choice.
System Requirements
Section titled “System Requirements”- CentOS Stream 9 / AlmaLinux 9 / Rocky Linux 9
- GCC 11+ or Clang 14+
- CMake 3.21+
- At least 4 GB RAM (for running 7B quantized models)
- Optional: NVIDIA GPU + CUDA (for GPU acceleration)
Install Build Dependencies
Section titled “Install Build Dependencies”sudo dnf install -y git gcc gcc-c++ make cmakeIf CUDA support is needed, make sure the NVIDIA driver and CUDA toolkit are installed (see NVIDIA Drivers and CUDA).
Build from Source
Section titled “Build from Source”Clone the Repository
Section titled “Clone the Repository”cd /optsudo mkdir -p llama-cpp && sudo chown $(whoami):$(whoami) llama-cppgit clone https://github.com/ggerganov/llama.cpp.git /opt/llama-cppcd /opt/llama-cppCPU-Only Build
Section titled “CPU-Only Build”cd /opt/llama-cppcmake -B buildcmake --build build --config Release -j$(nproc)After compilation, the binaries are located in the build/bin/ directory.
Build with CUDA GPU Acceleration
Section titled “Build with CUDA GPU Acceleration”cd /opt/llama-cppcmake -B build -DGGML_CUDA=ONcmake --build build --config Release -j$(nproc)Verify the CUDA build succeeded:
./build/bin/llama-cli --versionOther Build Options
Section titled “Other Build Options”# Enable OpenBLAS acceleration (CPU matrix operation optimization)sudo dnf install -y openblas-develcmake -B build -DGGML_BLAS=ON -DGGML_BLAS_VENDOR=OpenBLAScmake --build build --config Release -j$(nproc)
# Enable Vulkan (for AMD GPUs or NVIDIA GPUs without CUDA)sudo dnf install -y vulkan-headers vulkan-loader-develcmake -B build -DGGML_VULKAN=ONcmake --build build --config Release -j$(nproc)GGUF Model Format
Section titled “GGUF Model Format”GGUF (GPT-Generated Unified Format) is the model format used by llama.cpp. It supports various quantization levels, striking a balance between model size and inference quality.
Download GGUF Models
Section titled “Download GGUF Models”Download pre-quantized GGUF models from Hugging Face:
mkdir -p /opt/llama-cpp/models
# Download Qwen2.5-7B-Instruct GGUF (Q4_K_M quantization, approximately 4.7 GB)cd /opt/llama-cpp/modelswget https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-GGUF/resolve/main/qwen2.5-7b-instruct-q4_k_m.gguf
# Download Llama 3.1 8B GGUFwget https://huggingface.co/bartowski/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q4_K_M.ggufQuantization Levels
Section titled “Quantization Levels”| Quantization Type | Size (7B Model) | Quality | Speed |
|---|---|---|---|
| Q2_K | ~2.7 GB | Lower | Fastest |
| Q3_K_M | ~3.3 GB | Fair | Fast |
| Q4_K_M | ~4.7 GB | Good (Recommended) | Moderately fast |
| Q5_K_M | ~5.3 GB | Very good | Medium |
| Q6_K | ~5.9 GB | Excellent | Moderately slow |
| Q8_0 | ~7.2 GB | Near original | Slow |
| F16 | ~14 GB | Original precision | Slowest |
Q4_K_M is generally recommended as it provides a good balance between quality and size.
Running Inference
Section titled “Running Inference”Interactive Chat
Section titled “Interactive Chat”cd /opt/llama-cpp
./build/bin/llama-cli \ -m models/qwen2.5-7b-instruct-q4_k_m.gguf \ -c 4096 \ -n 512 \ --chat-template chatml \ -cnvParameter descriptions:
| Parameter | Description |
|---|---|
-m | Model file path |
-c | Context length |
-n | Maximum number of tokens to generate |
--chat-template | Chat template format |
-cnv | Enable conversation mode |
-t | Number of threads (auto-detected by default) |
-ngl | Number of layers to offload to GPU (available with CUDA builds) |
GPU Offloading
Section titled “GPU Offloading”If compiled with CUDA, you can offload some or all model layers to the GPU:
# Offload all layers to GPU (requires sufficient GPU VRAM)./build/bin/llama-cli \ -m models/qwen2.5-7b-instruct-q4_k_m.gguf \ -c 4096 \ -ngl 99 \ -cnv
# Offload only some layers (when VRAM is limited)./build/bin/llama-cli \ -m models/qwen2.5-7b-instruct-q4_k_m.gguf \ -c 4096 \ -ngl 20 \ -cnvSingle Generation
Section titled “Single Generation”./build/bin/llama-cli \ -m models/qwen2.5-7b-instruct-q4_k_m.gguf \ -p "Explain what a Linux kernel is in simple terms" \ -n 256 \ -c 2048Performance Benchmarking
Section titled “Performance Benchmarking”./build/bin/llama-bench \ -m models/qwen2.5-7b-instruct-q4_k_m.gguf \ -t $(nproc)Server Mode
Section titled “Server Mode”llama.cpp includes a built-in HTTP server that provides an OpenAI API-compatible interface.
Start the Server
Section titled “Start the Server”cd /opt/llama-cpp
./build/bin/llama-server \ -m models/qwen2.5-7b-instruct-q4_k_m.gguf \ --host 0.0.0.0 \ --port 8080 \ -c 4096 \ -ngl 99 \ --chat-template chatmlAPI Calls
Section titled “API Calls”The server provides an OpenAI-compatible API:
# Chat Completionscurl http://localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "qwen2.5", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is SELinux?"} ], "temperature": 0.7, "max_tokens": 512 }'# Text Completionscurl http://localhost:8080/v1/completions \ -H "Content-Type: application/json" \ -d '{ "prompt": "The advantages of the Linux operating system include", "max_tokens": 256 }'# Health checkcurl http://localhost:8080/healthThe server also includes a simple web interface accessible by visiting http://your-server-ip:8080 in your browser.
systemd Service
Section titled “systemd Service”sudo tee /etc/systemd/system/llama-server.service > /dev/null <<'EOF'[Unit]Description=llama.cpp ServerAfter=network-online.targetWants=network-online.target
[Service]Type=simpleUser=rootWorkingDirectory=/opt/llama-cppExecStart=/opt/llama-cpp/build/bin/llama-server \ -m /opt/llama-cpp/models/qwen2.5-7b-instruct-q4_k_m.gguf \ --host 0.0.0.0 \ --port 8080 \ -c 4096 \ -ngl 99 \ --chat-template chatmlRestart=on-failureRestartSec=10LimitNOFILE=65536
[Install]WantedBy=multi-user.targetEOF
sudo systemctl daemon-reloadsudo systemctl enable --now llama-serverModel Quantization
Section titled “Model Quantization”If you have a raw Hugging Face model, you can convert and quantize it yourself using llama.cpp.
Convert a Model to GGUF
Section titled “Convert a Model to GGUF”cd /opt/llama-cpp
# Install Python dependencies (only needed for conversion)pip install -r requirements.txt
# Convert a Hugging Face model to GGUF (F16 precision)python convert_hf_to_gguf.py /path/to/huggingface-model/ \ --outfile models/my-model-f16.gguf \ --outtype f16Quantize a Model
Section titled “Quantize a Model”# Quantize an F16 model to Q4_K_M./build/bin/llama-quantize \ models/my-model-f16.gguf \ models/my-model-q4_k_m.gguf \ Q4_K_M
# Quantize to Q5_K_M (higher quality)./build/bin/llama-quantize \ models/my-model-f16.gguf \ models/my-model-q5_k_m.gguf \ Q5_K_MVerify a Quantized Model
Section titled “Verify a Quantized Model”# Quick test to verify the quantized model works./build/bin/llama-cli \ -m models/my-model-q4_k_m.gguf \ -p "Hello, how are you?" \ -n 64Firewall Configuration
Section titled “Firewall Configuration”sudo firewall-cmd --permanent --add-port=8080/tcpsudo firewall-cmd --reloadUpdating llama.cpp
Section titled “Updating llama.cpp”llama.cpp is updated very frequently. It is recommended to update regularly for performance improvements and new model support:
cd /opt/llama-cppgit pull
# Clean old build files and recompilerm -rf buildcmake -B build -DGGML_CUDA=ONcmake --build build --config Release -j$(nproc)
# Restart the servicesudo systemctl restart llama-serverBuild Error: CMake Version Too Old
Section titled “Build Error: CMake Version Too Old”# Install a newer version of CMakesudo dnf install -y cmake3# Or compile CMake from sourceSlow Inference Speed
Section titled “Slow Inference Speed”# Check if all CPU cores are being used./build/bin/llama-cli -m model.gguf -t $(nproc) ...
# If you have a GPU, make sure you compiled with CUDA and set -ngl./build/bin/llama-cli -m model.gguf -ngl 99 ...
# Use a smaller quantized model# Q4_K_M is typically 40-50% faster than Q8_0Out of Memory
Section titled “Out of Memory”# Use a smaller quantization level# Q2_K is about 1/5 the size of F16
# Reduce context length./build/bin/llama-cli -m model.gguf -c 2048 ...
# Use mmap (enabled by default) to avoid loading the entire model into memoryCUDA-Related Errors
Section titled “CUDA-Related Errors”# Verify CUDA toolkit versionnvcc --version
# Verify CUDA was enabled during compilationcmake -B build -DGGML_CUDA=ONcmake --build build --config Release -j$(nproc)
# Check if the GPU is detectednvidia-smi