Ollama Local LLM
Ollama is a lightweight local large language model runtime framework that lets you easily run open-source models like Llama 3, Qwen2, and Gemma on your own server without relying on cloud APIs. It provides a clean command-line tool and an OpenAI-compatible REST API, making it ideal for deploying private AI services on enterprise intranets.
System Requirements
Section titled “System Requirements”- CentOS Stream 9 / AlmaLinux 9 / Rocky Linux 9
- At least 8 GB RAM (for running 7B models); 16 GB or more recommended
- For GPU acceleration: NVIDIA GPU with drivers and CUDA installed (see NVIDIA Drivers and CUDA)
- CPU mode is also supported but slower
Installing Ollama
Section titled “Installing Ollama”Ollama provides a one-line installation script that works on all major Linux distributions:
curl -fsSL https://ollama.com/install.sh | shAfter installation, Ollama will automatically register as a systemd service and start. Verify the installation:
ollama --versionCheck the service status:
sudo systemctl status ollamaManual Installation (Offline Environments)
Section titled “Manual Installation (Offline Environments)”If the server cannot access the internet, you can install manually:
# Download the binary for your architecturecurl -L https://ollama.com/download/ollama-linux-amd64.tgz -o ollama-linux-amd64.tgz
# Extract to /usr directorysudo tar -C /usr -xzf ollama-linux-amd64.tgz
# Verifyollama --versionCreate a dedicated user and group:
sudo useradd -r -s /bin/false -U -m -d /usr/share/ollama ollamasudo usermod -a -G ollama $(whoami)Pulling Models
Section titled “Pulling Models”Ollama manages models in a Docker-like fashion. Use ollama pull to download models.
# Pull Meta Llama 3.1 8B (approximately 4.7 GB)ollama pull llama3.1
# Pull Alibaba Qwen2.5 7B (approximately 4.7 GB)ollama pull qwen2.5
# Pull Google Gemma 2 9B (approximately 5.4 GB)ollama pull gemma2
# Pull smaller models suitable for low-spec machinesollama pull llama3.2:3bollama pull qwen2.5:1.5bList downloaded models:
ollama listDelete a model you no longer need:
ollama rm llama3.1Running Models
Section titled “Running Models”Interactive Chat
Section titled “Interactive Chat”# Start an interactive chat sessionollama run llama3.1Once in the chat, type your questions directly. Type /bye to exit.
Single Invocation
Section titled “Single Invocation”# Pipe a question inecho "Explain what Linux is in one sentence" | ollama run qwen2.5Multi-line Input
Section titled “Multi-line Input”In interactive mode, use """ to begin and end multi-line input:
>>> """Please translate the following into English:Ollama is a powerful local large model framework."""API Usage
Section titled “API Usage”Ollama provides a REST API service at http://localhost:11434 by default.
Generate Text (Generate)
Section titled “Generate Text (Generate)”curl http://localhost:11434/api/generate -d '{ "model": "qwen2.5", "prompt": "Please explain what SELinux is in Linux", "stream": false}'Chat Endpoint (Chat)
Section titled “Chat Endpoint (Chat)”curl http://localhost:11434/api/chat -d '{ "model": "qwen2.5", "messages": [ { "role": "system", "content": "You are a Linux system administration expert." }, { "role": "user", "content": "How do I check the CentOS system version?" } ], "stream": false}'OpenAI-Compatible Endpoint
Section titled “OpenAI-Compatible Endpoint”Ollama also provides an endpoint compatible with the OpenAI API format, making it easy to integrate with existing tools:
curl http://localhost:11434/v1/chat/completions -d '{ "model": "qwen2.5", "messages": [ {"role": "user", "content": "Hello"} ]}'This means you can use Ollama as a drop-in replacement for the OpenAI API backend by simply pointing your API URL to http://localhost:11434/v1.
List Available Models
Section titled “List Available Models”curl http://localhost:11434/api/tagssystemd Service Management
Section titled “systemd Service Management”The installation script automatically creates a systemd service. If you installed manually, you need to create the service file yourself:
sudo tee /etc/systemd/system/ollama.service > /dev/null <<'EOF'[Unit]Description=Ollama ServiceAfter=network-online.target
[Service]ExecStart=/usr/bin/ollama serveUser=ollamaGroup=ollamaRestart=alwaysRestartSec=3Environment="HOME=/usr/share/ollama"Environment="OLLAMA_HOST=0.0.0.0"
[Install]WantedBy=default.targetEOFEnable and start the service:
sudo systemctl daemon-reloadsudo systemctl enable --now ollamaCommon service management commands:
# Check statussudo systemctl status ollama
# View logssudo journalctl -u ollama -f
# Restart servicesudo systemctl restart ollama
# Stop servicesudo systemctl stop ollamaEnvironment Variable Configuration
Section titled “Environment Variable Configuration”Modify configuration through the systemd override mechanism:
sudo systemctl edit ollamaAdd the desired environment variables in the editor:
[Service]# Listen on all network interfaces (default is localhost only)Environment="OLLAMA_HOST=0.0.0.0"# Change model storage directoryEnvironment="OLLAMA_MODELS=/data/ollama/models"# Set number of concurrent requestsEnvironment="OLLAMA_NUM_PARALLEL=4"# Set maximum number of loaded modelsEnvironment="OLLAMA_MAX_LOADED_MODELS=2"After making changes, reload and restart:
sudo systemctl daemon-reloadsudo systemctl restart ollamaGPU Acceleration (NVIDIA CUDA)
Section titled “GPU Acceleration (NVIDIA CUDA)”Ollama automatically detects NVIDIA GPUs and uses CUDA acceleration, provided that NVIDIA drivers are properly installed.
Confirm GPU Is Detected
Section titled “Confirm GPU Is Detected”# Check NVIDIA drivernvidia-smi
# Check if Ollama detects the GPUollama psAfter running a model, you can see GPU memory being used via nvidia-smi:
# Run a model firstollama run llama3.1 "hello" && nvidia-smiSpecify GPU
Section titled “Specify GPU”If you have multiple GPUs, you can specify which ones to use via environment variables:
sudo systemctl edit ollama[Service]# Use only GPU 0 and GPU 1Environment="CUDA_VISIBLE_DEVICES=0,1"CPU Only
Section titled “CPU Only”If you have a GPU but want to force CPU-only execution:
CUDA_VISIBLE_DEVICES="" ollama run llama3.1Custom Models (Modelfile)
Section titled “Custom Models (Modelfile)”You can create custom model configurations using a Modelfile:
cat > ~/Modelfile <<'EOF'FROM qwen2.5
SYSTEM """You are an experienced Linux system administrator who specializes in CentOS/AlmaLinux/Rocky Linux server operations.When answering questions, provide specific commands and steps."""
PARAMETER temperature 0.3PARAMETER num_ctx 4096EOF
# Create the custom modelollama create linux-assistant -f ~/Modelfile
# Run the custom modelollama run linux-assistantFirewall Configuration
Section titled “Firewall Configuration”If you need to access the Ollama API from other machines, open the port:
sudo firewall-cmd --permanent --add-port=11434/tcpsudo firewall-cmd --reloadSlow Model Downloads
Section titled “Slow Model Downloads”You can set a proxy to speed up downloads:
sudo systemctl edit ollama[Service]Environment="HTTPS_PROXY=http://your-proxy:7890"Environment="HTTP_PROXY=http://your-proxy:7890"Environment="NO_PROXY=localhost,127.0.0.1"Out of Memory
Section titled “Out of Memory”If you encounter OOM (Out of Memory) errors, try using smaller quantized models:
# Use 4-bit quantized small modelsollama pull llama3.2:3bollama pull qwen2.5:1.5bView Model Details
Section titled “View Model Details”ollama show llama3.1ollama show llama3.1 --modelfile