The Coe Lab
← Back to Blog

How to Self-Host Ollama with Docker: Complete 2026 Guide

September 28, 20268 min read
OllamaDockerself-hostingLLMhomelab

Step-by-step guide to self-hosting Ollama with Docker in 2026. Install, configure GPU passthrough, expose APIs, and run local LLMs securely on your own hardware.

Running large language models locally has never been easier. Ollama, combined with Docker, gives you a clean, reproducible way to serve LLMs on your own hardware — whether that's a beefy homelab server, a cloud VPS, or a workstation with a spare GPU. This guide walks through the complete setup from scratch, including GPU passthrough, API exposure, and production hardening.

Why Self-Host Ollama?

Ollama has become the de facto standard for running LLMs locally. It handles model downloading, quantization, and serving through a clean REST API. By self-hosting, you get complete data privacy, no per-token API costs, and the ability to fine-tune or customize models for your workloads. Docker makes the deployment reproducible and easy to update.

The main use cases for self-hosted Ollama include: building AI-powered applications without API costs, running private coding assistants, powering RAG pipelines with local embeddings, and experimenting with open-weight models like Llama, Qwen, and Mistral.

Prerequisites

Before starting, make sure you have:

  • A machine running Linux (Ubuntu 22.04+ recommended) or macOS
  • Docker and Docker Compose installed
  • At least 8GB RAM (16GB+ recommended for larger models)
  • An NVIDIA GPU with at least 8GB VRAM (optional but strongly recommended for performance)
  • NVIDIA Container Toolkit if using GPU passthrough
  • 20GB+ free disk space for model storage

Step 1: Install Docker and Docker Compose

If Docker isn't already installed, start with the official installation script. On Ubuntu or Debian-based systems:

curl -fsSL https://get.docker.com | sudo sh # Add your user to the docker group sudo usermod -aG docker $USER # Verify installation docker --version docker compose version

Log out and back in for the group change to take effect, or run `newgrp docker` in your current terminal.

Step 2: Set Up GPU Passthrough (NVIDIA)

If you have an NVIDIA GPU, install the NVIDIA Container Toolkit so Docker containers can access it. Skip this section if you're running CPU-only.

# Add NVIDIA's package repository distribution=$(. /etc/os-release; echo $ID$VERSION_ID) curl -s -L https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list | \ sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \ sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list sudo apt-get update sudo apt-get install -y nvidia-container-toolkit # Configure Docker to use NVIDIA runtime sudo nvidia-ctk runtime configure --runtime=docker sudo systemctl restart docker

Verify GPU access from Docker:

docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi

You should see your GPU's output. If not, check that your NVIDIA drivers are installed (`nvidia-smi` works on the host) and that the container toolkit is properly configured.

Step 3: Create the Docker Compose File

Create a project directory and a docker-compose.yml file. This configuration persists model data and exposes the Ollama API on port 11434:

mkdir -p ~/ollama && cd ~/ollama cat > docker-compose.yml << 'EOF' services: ollama: image: ollama/ollama:latest container_name: ollama restart: unless-stopped ports: - "11434:11434" volumes: - ollama_data:/root/.ollama # For GPU support, uncomment these lines: # deploy: # resources: # reservations: # devices: # - driver: nvidia # count: all # capabilities: [gpu] volumes: ollama_data: EOF

If you set up GPU passthrough in Step 2, uncomment the deploy section. The volume ensures your downloaded models persist across container restarts and updates.

Step 4: Start Ollama

Pull and start the container:

docker compose up -d # Check logs docker compose logs -f ollama

Ollama should start and show that it's listening on port 11434. Verify it's running:

curl http://localhost:11434/api/version

You should get a JSON response with the version number. If the connection fails, check the Docker logs for errors.

Step 5: Download and Run Your First Model

Ollama makes it trivial to download and run models. The command below pulls a compact but capable model and starts a chat session:

# Pull a model (Qwen 2.5 7B is a great starting point) docker exec -it ollama ollama pull qwen2.5:7b # Start a chat session docker exec -it ollama ollama run qwen2.5:7b

Some popular models to try:

  • qwen2.5:7b — Fast, capable, great for general tasks (4.4GB download)
  • llama3.2:8b — Meta's latest mid-size model (4.7GB download)
  • mistral:7b — Mistral's efficient model, excellent for coding (4.1GB download)
  • qwen2.5-coder:7b — Specialized for code generation and completion
  • nomic-embed-text — For generating text embeddings (RAG pipelines)

Model size matters. A 7B parameter model needs roughly 5-6GB of RAM or VRAM in Q4 quantization. If you have a GPU with 12GB+ VRAM, you can run 13B models. With 24GB VRAM, you can handle 33B models. Check your available memory with `nvidia-smi` before pulling large models.

Step 6: Use the Ollama REST API

Ollama exposes a clean REST API that you can use from any application. Here are the key endpoints:

Generate a completion:

curl http://localhost:11434/api/generate -d '{ "model": "qwen2.5:7b", "prompt": "Write a Python function to reverse a linked list", "stream": false }'

Chat with conversation history:

curl http://localhost:11434/api/chat -d '{ "model": "qwen2.5:7b", "messages": [ {"role": "user", "content": "Explain Docker volumes in simple terms"} ], "stream": false }'

List installed models:

curl http://localhost:11434/api/tags

The API supports streaming responses (set `"stream": true`), temperature control, system prompts, and JSON mode output. For production applications, wrap these calls in your preferred language's HTTP client.

Step 7: Production Hardening

For a production deployment, you'll want to add several layers of security and reliability:

Reverse Proxy with TLS

Don't expose Ollama directly to the internet. Put it behind a reverse proxy like Caddy or Nginx with TLS termination:

# Add Caddy to your docker-compose.yml caddy: image: caddy:latest container_name: caddy restart: unless-stopped ports: - "80:80" - "443:443" volumes: - ./Caddyfile:/etc/caddy/Caddyfile - caddy_data:/data depends_on: - ollama

Create a Caddyfile with your domain:

ollama.yourdomain.com { reverse_proxy ollama:11434 # Add basic auth basicauth { admin $2a$14$yourhashedpassword } }

API Key Authentication

Ollama doesn't have built-in authentication. Use your reverse proxy to add API key checks or basic auth. For internal use on a homelab network, consider a WireGuard tunnel instead of exposing the API publicly.

Automatic Backups

Back up your models directory regularly. Since models are stored in a Docker volume, you can export them:

# Backup all models docker run --rm -v ollama_ollama_data:/data -v $(pwd):/backup alpine \ tar czf /backup/ollama-backup-$(date +%Y%m%d).tar.gz /data # Restore from backup docker run --rm -v ollama_ollama_data:/data -v $(pwd):/backup alpine \ tar xzf /backup/ollama-backup-20260928.tar.gz -C /

Step 8: Building Applications on Top of Ollama

Once Ollama is running, you can build applications that use it. Popular options include:

  • Open WebUI — A ChatGPT-like web interface that connects to Ollama
  • Continue.dev — VS Code extension for AI-powered coding with local models
  • LangChain / LlamaIndex — Python frameworks for building RAG pipelines and agents
  • AnythingLLM — Desktop app for document chat with local LLMs

To add Open WebUI to your Docker Compose setup:

open-webui: image: ghcr.io/open-webui/open-webui:main container_name: open-webui restart: unless-stopped ports: - "3000:8080" environment: - OLLAMA_BASE_URL=http://ollama:11434 volumes: - open_webui_data:/app/backend/data depends_on: - ollama

Access Open WebUI at http://localhost:3000 and you'll have a full chat interface with model selection, conversation history, and document upload.

Troubleshooting Common Issues

Out of Memory Errors

If you see 'CUDA out of memory' or the model fails to load, you're trying to run a model too large for your available VRAM. Switch to a smaller quantization or a smaller model. You can also set `OLLAMA_MAX_VRAM` environment variable to limit GPU memory usage.

Slow Inference on CPU

CPU inference is 10-50x slower than GPU. If you don't have a GPU, consider using smaller models (1.5B or 3B parameters) or renting a cloud GPU instance. Ollama Cloud and various GPU cloud providers offer affordable hourly rates.

Model Download Failures

Large model downloads can timeout. Use `ollama pull` with the `--verbose` flag to see progress. If downloads fail repeatedly, check your network connection and disk space. You can also manually download models from Hugging Face and import them.

Docker Volume Permission Issues

If Ollama can't write to its data directory, check the volume permissions. On SELinux systems, you may need to add `:Z` to the volume mount: `-v ollama_data:/root/.ollama:Z`.

Cost Comparison: Self-Hosted vs Cloud API

The economics of self-hosting depend on your usage. A used RTX 3090 (24GB VRAM) costs around $600-700 and can run models up to 33B parameters. At that price point:

  • Cloud API cost for GPT-4-class model: ~$0.03-0.06 per 1K tokens
  • Self-hosted cost: electricity (~$10-20/month for a single GPU server)
  • Break-even point: roughly 500K-1M tokens per month
  • For high-volume workloads (coding assistants, RAG pipelines), self-hosting pays for itself in 1-3 months

The tradeoff is maintenance. Cloud APIs handle updates, scaling, and uptime. With self-hosting, you're responsible for keeping your system updated, monitoring GPU health, and handling failures.

Conclusion

Self-hosting Ollama with Docker gives you a production-grade local LLM platform in under an hour. The combination of Docker's reproducibility and Ollama's simplicity makes it accessible even if you're not a DevOps expert. Start with a 7B model on a single GPU, add Open WebUI for a chat interface, and scale up as your needs grow.

The key advantages — data privacy, no per-token costs, and full control over your AI infrastructure — make self-hosting increasingly attractive as open-weight models approach the quality of proprietary alternatives. Whether you're building AI-powered applications, running a private coding assistant, or just experimenting with local LLMs, this setup gives you a solid foundation to build on.

Related Posts

Strata: When a 125B AI Model Ran on a Gaming PC at 100 Tokens Per Second

A new open-source tool called Strata lets you run Qwen 3.8 Flash Next — a 125-billion-parameter model — on an ordinary gaming PC with an RTX 4090. Nothing leaves your machine, and it's faster than you can read.

Oct 5, 2026• 6 min

When AI Agents Spend Your Money While You Sleep: Why Hard Budget Caps Are Becoming Non-Negotiable

AWS and Google Cloud finally launched hard spending limits in the same month. It's not a coincidence — it's a response to AI agents that can rack up thousands of dollars before you wake up.

Oct 4, 2026• 6 min

When Utah Banned VPNs: How a Court Stopped a Law That Demanded the Technically Impossible

A federal judge just blocked Utah's unprecedented anti-VPN law, ruling that lawmakers cannot mandate perfect geolocation — a technical impossibility. The case reveals a deeper problem: when legislation outruns engineering.

Oct 3, 2026• 7 min