Skip to content

Ollama Local LLM

Ollama is a lightweight local large language model runtime framework that lets you easily run open-source models like Llama 3, Qwen2, and Gemma on your own server without relying on cloud APIs. It provides a clean command-line tool and an OpenAI-compatible REST API, making it ideal for deploying private AI services on enterprise intranets.

  • CentOS Stream 9 / AlmaLinux 9 / Rocky Linux 9
  • At least 8 GB RAM (for running 7B models); 16 GB or more recommended
  • For GPU acceleration: NVIDIA GPU with drivers and CUDA installed (see NVIDIA Drivers and CUDA)
  • CPU mode is also supported but slower

Ollama provides a one-line installation script that works on all major Linux distributions:

Terminal window
curl -fsSL https://ollama.com/install.sh | sh

After installation, Ollama will automatically register as a systemd service and start. Verify the installation:

Terminal window
ollama --version

Check the service status:

Terminal window
sudo systemctl status ollama

Manual Installation (Offline Environments)

Section titled “Manual Installation (Offline Environments)”

If the server cannot access the internet, you can install manually:

Terminal window
# Download the binary for your architecture
curl -L https://ollama.com/download/ollama-linux-amd64.tgz -o ollama-linux-amd64.tgz
# Extract to /usr directory
sudo tar -C /usr -xzf ollama-linux-amd64.tgz
# Verify
ollama --version

Create a dedicated user and group:

Terminal window
sudo useradd -r -s /bin/false -U -m -d /usr/share/ollama ollama
sudo usermod -a -G ollama $(whoami)

Ollama manages models in a Docker-like fashion. Use ollama pull to download models.

Terminal window
# Pull Meta Llama 3.1 8B (approximately 4.7 GB)
ollama pull llama3.1
# Pull Alibaba Qwen2.5 7B (approximately 4.7 GB)
ollama pull qwen2.5
# Pull Google Gemma 2 9B (approximately 5.4 GB)
ollama pull gemma2
# Pull smaller models suitable for low-spec machines
ollama pull llama3.2:3b
ollama pull qwen2.5:1.5b

List downloaded models:

Terminal window
ollama list

Delete a model you no longer need:

Terminal window
ollama rm llama3.1
Terminal window
# Start an interactive chat session
ollama run llama3.1

Once in the chat, type your questions directly. Type /bye to exit.

Terminal window
# Pipe a question in
echo "Explain what Linux is in one sentence" | ollama run qwen2.5

In interactive mode, use """ to begin and end multi-line input:

>>> """
Please translate the following into English:
Ollama is a powerful local large model framework.
"""

Ollama provides a REST API service at http://localhost:11434 by default.

Terminal window
curl http://localhost:11434/api/generate -d '{
"model": "qwen2.5",
"prompt": "Please explain what SELinux is in Linux",
"stream": false
}'
Terminal window
curl http://localhost:11434/api/chat -d '{
"model": "qwen2.5",
"messages": [
{
"role": "system",
"content": "You are a Linux system administration expert."
},
{
"role": "user",
"content": "How do I check the CentOS system version?"
}
],
"stream": false
}'

Ollama also provides an endpoint compatible with the OpenAI API format, making it easy to integrate with existing tools:

Terminal window
curl http://localhost:11434/v1/chat/completions -d '{
"model": "qwen2.5",
"messages": [
{"role": "user", "content": "Hello"}
]
}'

This means you can use Ollama as a drop-in replacement for the OpenAI API backend by simply pointing your API URL to http://localhost:11434/v1.

Terminal window
curl http://localhost:11434/api/tags

The installation script automatically creates a systemd service. If you installed manually, you need to create the service file yourself:

Terminal window
sudo tee /etc/systemd/system/ollama.service > /dev/null <<'EOF'
[Unit]
Description=Ollama Service
After=network-online.target
[Service]
ExecStart=/usr/bin/ollama serve
User=ollama
Group=ollama
Restart=always
RestartSec=3
Environment="HOME=/usr/share/ollama"
Environment="OLLAMA_HOST=0.0.0.0"
[Install]
WantedBy=default.target
EOF

Enable and start the service:

Terminal window
sudo systemctl daemon-reload
sudo systemctl enable --now ollama

Common service management commands:

Terminal window
# Check status
sudo systemctl status ollama
# View logs
sudo journalctl -u ollama -f
# Restart service
sudo systemctl restart ollama
# Stop service
sudo systemctl stop ollama

Modify configuration through the systemd override mechanism:

Terminal window
sudo systemctl edit ollama

Add the desired environment variables in the editor:

[Service]
# Listen on all network interfaces (default is localhost only)
Environment="OLLAMA_HOST=0.0.0.0"
# Change model storage directory
Environment="OLLAMA_MODELS=/data/ollama/models"
# Set number of concurrent requests
Environment="OLLAMA_NUM_PARALLEL=4"
# Set maximum number of loaded models
Environment="OLLAMA_MAX_LOADED_MODELS=2"

After making changes, reload and restart:

Terminal window
sudo systemctl daemon-reload
sudo systemctl restart ollama

Ollama automatically detects NVIDIA GPUs and uses CUDA acceleration, provided that NVIDIA drivers are properly installed.

Terminal window
# Check NVIDIA driver
nvidia-smi
# Check if Ollama detects the GPU
ollama ps

After running a model, you can see GPU memory being used via nvidia-smi:

Terminal window
# Run a model first
ollama run llama3.1 "hello" && nvidia-smi

If you have multiple GPUs, you can specify which ones to use via environment variables:

Terminal window
sudo systemctl edit ollama
[Service]
# Use only GPU 0 and GPU 1
Environment="CUDA_VISIBLE_DEVICES=0,1"

If you have a GPU but want to force CPU-only execution:

Terminal window
CUDA_VISIBLE_DEVICES="" ollama run llama3.1

You can create custom model configurations using a Modelfile:

Terminal window
cat > ~/Modelfile <<'EOF'
FROM qwen2.5
SYSTEM """
You are an experienced Linux system administrator who specializes in CentOS/AlmaLinux/Rocky Linux server operations.
When answering questions, provide specific commands and steps.
"""
PARAMETER temperature 0.3
PARAMETER num_ctx 4096
EOF
# Create the custom model
ollama create linux-assistant -f ~/Modelfile
# Run the custom model
ollama run linux-assistant

If you need to access the Ollama API from other machines, open the port:

Terminal window
sudo firewall-cmd --permanent --add-port=11434/tcp
sudo firewall-cmd --reload

You can set a proxy to speed up downloads:

Terminal window
sudo systemctl edit ollama
[Service]
Environment="HTTPS_PROXY=http://your-proxy:7890"
Environment="HTTP_PROXY=http://your-proxy:7890"
Environment="NO_PROXY=localhost,127.0.0.1"

If you encounter OOM (Out of Memory) errors, try using smaller quantized models:

Terminal window
# Use 4-bit quantized small models
ollama pull llama3.2:3b
ollama pull qwen2.5:1.5b
Terminal window
ollama show llama3.1
ollama show llama3.1 --modelfile