Self-hosted S3-compatible object storage for backups: MinIO behind an Nginx proxy, with an rclone-based backup agent and versioned buckets. Your offsite backups without a cloud bill.
A complete Kubernetes lab-to-production pattern: k3s in Docker with the embedded etcd datastore, Rancher for management, MetalLB for Layer-2 LoadBalancer IPs, Longhorn replicated storage, and Traefik
Run LLMs on your own GPU hardware: Ollama as the serving engine, Open WebUI as the ChatGPT-style frontend, and LiteLLM as an OpenAI-compatible router in front of everything — so local models and external APIs stay interchangeable.
A deployment guide with working examples — not a downloadable product. Everything below is the actual configuration I use, abridged to the parts that matter.
services:
ollama:
image: ollama/ollama:${OLLAMA_VERSION:-0.5.9}
container_name: ollama
restart: unless-stopped
volumes:
- ollama_models:/root/.ollama
- ./config/ollama:/etc/ollama:ro
ports:
- ${OLLAMA_PORT:-11434}:11434
environment:
- OLLAMA_HOST=0.0.0.0
- OLLAMA_KEEP_ALIVE=${OLLAMA_KEEP_ALIVE:-5m}
- OLLAMA_NUM_PARALLEL=${OLLAMA_NUM_PARALLEL:-2}
- OLLAMA_MAX_LOADED_MODELS=${OLLAMA_MAX_LOADED_MODELS:-3}
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities:
- gpu
healthcheck:
test:
- CMD
- curl
- -f
- http://localhost:11434/api/tags
interval: 30s
timeout: 10s
retries: 3
networks:
- ai_network
litellm:
image: ghcr.io/berriai/litellm:${LITELLM_VERSION:-main-v1.16.18}
container_name: litellm
restart: unless-stopped
depends_on:
ollama:
condition: service_healthy
volumes:
- ./config/litellm/proxy_config.yaml:/app/proxy_config.yaml:ro
ports:
- ${LITELLM_PORT:-4000}:4000
environment:
- LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY:?LITELLM_MASTER_KEY is required}
- DATABASE_URL=${LITELLM_DB_URL:-}
- LITELLM_MODE=${LITELLM_MODE:-proxy}
- OPENAI_API_KEY=${OPENAI_API_KEY:-}
command:
- --config
- /app/proxy_config.yaml
- --port
- '4000'
healthcheck:
test:
- CMD
- curl
- -f
- http://localhost:4000/health/liveliness
interval: 30s
timeout: 10s
retries: 3
networks:
- ai_network
Excerpt — ollama, litellm from the compose file. The remaining services (proxies, init jobs, exporters) follow the same pattern and mount their configuration from a config/ directory.
Copy .env.example to .env and at minimum set:
OLLAMA_VERSION=latest
OLLAMA_PORT=11434
OLLAMA_HOST=0.0.0.0
OLLAMA_KEEP_ALIVE=10m
OLLAMA_NUM_PARALLEL=2
OLLAMA_MAX_LOADED_MODELS=4
.env before the first start — never ship the example valuesA guide, not a product. This page is deployment documentation with working examples — there is no zip, no download, no support contract. You adapt the patterns to your own environment, and you own the result.
Want this running production-grade in your infrastructure instead? This is the core of the two-day Self-Hosted AI in Production workshop — details on the consulting page.