Metrics tell you a service is dying. Logs tell you why. If you’ve ever SSH’d into three different containers in a row trying to figure out which one actually threw the error that broke your reverse proxy, you already know the gap that centralized logging fills. Uptime checks and Grafana dashboards (covered in the monitoring stack guide) tell you something is wrong and roughly when. Logs tell you what the service was actually doing in the seconds before it went wrong, and without them you’re stuck doing docker logs on one container at a time, hoping you guessed right on the first try.

The homelab version of this problem is smaller than what these tools were built for, but the pain is identical: logs scattered across a dozen containers and a couple of hosts, no shared timeline, and no way to search across all of them at once when something breaks at 2am and you need an answer fast.

What centralized logging actually buys you

Three things, in order of how often you’ll actually use them:

  1. One search box across everything. Grep a hostname, a request ID, or an error string and get results from every container and host that logged it, in time order, instead of opening a terminal per service.
  2. Correlation across services. A failed login on your auth service and a 500 from the app behind it, in the same timeline, make root cause obvious instead of requiring you to mentally merge two separate log streams.
  3. Retention past what the container runtime keeps. Docker’s default logging driver rotates logs and throws old ones away. Once a container restarts, whatever was in its logs before that is usually gone unless you shipped it somewhere else first.

None of this prevents an outage. It cuts the time between “something’s wrong” and “I know what broke it” from twenty minutes of log-spelunking to a two-minute search.

The three real options, and who they’re actually for

Grafana Loki is the lightest of the three and the one most homelabs should start with, especially if Grafana is already running for metrics. Loki doesn’t index log content the way the other two do, it indexes labels (container name, host, job) and stores the raw log lines compressed alongside them. That tradeoff means much lower resource usage and storage cost, at the price of slower full-text search across huge volumes - a tradeoff that’s irrelevant at homelab scale, where you’re searching gigabytes, not terabytes, per day.

Graylog sits in the middle. It’s built on Elasticsearch/OpenSearch under the hood but wraps it in a purpose-built log-management UI with saved searches, dashboards, and alerting rules that are easier to reach for than wiring up Kibana from scratch. It wants more RAM than Loki (realistically 4GB+ for the stack, more if retention grows), but the UI is more approachable for someone who wants a real log-management tool rather than a Grafana plugin.

The ELK/Elastic stack (Elasticsearch, Logstash, Kibana, or OpenSearch’s fork of the same shape) is the heaviest and most capable option, and genuinely more than a single-operator homelab needs unless you’re specifically trying to learn it for work. Elasticsearch alone wants multiple gigabytes of heap just to idle comfortably, Logstash’s pipeline configuration has a real learning curve, and the thing you get at the end is enterprise-grade full-text search and analytics that a homelab’s log volume will never come close to stressing. If you’re running ELK at work and want the same tool at home for muscle memory, that’s a legitimate reason. If you just want to find out why a container crashed, it’s the wrong tool for the job size.

For almost every homelab, the real choice is Loki if Grafana’s already in the stack, Graylog if you want a dedicated log UI without touching Elasticsearch configuration directly.

Setting up Loki (the path most homelabs should take)

If Prometheus and Grafana are already running per the monitoring guide, Loki slots in as a third piece feeding the same Grafana frontend you’re already using. The missing piece is Promtail (or its newer replacement, Alloy), the agent that tails container logs and ships them to Loki with the right labels attached.

services:
  loki:
    image: grafana/loki:latest
    volumes:
      - ./loki-data:/loki
    ports:
      - "3100:3100"
    restart: unless-stopped

  promtail:
    image: grafana/promtail:latest
    volumes:
      - /var/lib/docker/containers:/var/lib/docker/containers:ro
      - /var/run/docker.sock:/var/run/docker.sock:ro
      - ./promtail-config.yml:/etc/promtail/config.yml
    command: -config.file=/etc/promtail/config.yml
    restart: unless-stopped

A minimal promtail-config.yml using Docker service discovery, so every container gets picked up automatically instead of hand-listing them:

server:
  http_listen_port: 9080

positions:
  filename: /tmp/positions.yaml

clients:
  - url: http://loki:3100/loki/api/v1/push

scrape_configs:
  - job_name: docker
    docker_sd_configs:
      - host: unix:///var/run/docker.sock
        refresh_interval: 15s
    relabel_configs:
      - source_labels: ['__meta_docker_container_name']
        target_label: 'container'

Add Loki as a data source in Grafana (http://loki:3100) and query it through the Explore tab using LogQL, Loki’s query language. It reads a lot like a filtered grep: {container="vaultwarden"} |= "error" pulls every log line from that container containing “error.” Pin a saved query for the services you check most often, same as you’d pin a dashboard.

If you’re running this inside an LXC rather than a VM, the same Docker-in-LXC networking gotcha from the monitoring guide applies here too - network_mode: host is the standard fix if containers can’t reach each other over the default bridge.

Setting up Graylog, if you want the dedicated UI instead

Graylog needs MongoDB (for configuration) and OpenSearch (for log storage) alongside itself, so it’s a heavier compose stack than Loki out of the gate:

services:
  mongodb:
    image: mongo:6
    volumes:
      - ./mongo-data:/data/db
    restart: unless-stopped

  opensearch:
    image: opensearchproject/opensearch:2
    environment:
      - "discovery.type=single-node"
      - "OPENSEARCH_JAVA_OPTS=-Xms1g -Xmx1g"
      - "DISABLE_SECURITY_PLUGIN=true"
    volumes:
      - ./opensearch-data:/usr/share/opensearch/data
    restart: unless-stopped

  graylog:
    image: graylog/graylog:6.0
    environment:
      - GRAYLOG_PASSWORD_SECRET=change-this-to-a-long-random-string
      - GRAYLOG_ROOT_PASSWORD_SHA2=sha256-hash-of-your-admin-password
      - GRAYLOG_HTTP_EXTERNAL_URI=http://graylog.yourdomain.lan/
    depends_on:
      - mongodb
      - opensearch
    ports:
      - "9000:9000"
      - "5140:5140/udp"
    restart: unless-stopped

GRAYLOG_ROOT_PASSWORD_SHA2 wants the SHA-256 hash of your chosen password, not the password itself - generate it with echo -n 'yourpassword' | sha256sum and paste the hash in.

Once it’s up, the web UI walks you through creating an Input - the listener that actually receives logs. GELF UDP is the lightest option for Docker containers; point your containers’ logging driver at it, or run a small sidecar like gelf-docker that forwards from the Docker socket the way Promtail does for Loki. From there, Graylog’s own search UI (closer to a purpose-built log tool than Grafana’s Explore tab) handles the querying, saved searches, and alert rules.

Making Docker actually ship its logs

Neither tool does anything until containers are configured to send logs somewhere other than Docker’s default json-file driver sitting on local disk. Two paths:

  • Sidecar/agent pattern (what both examples above use): leave each container’s logging driver alone and let Promtail or a GELF forwarder read from the Docker socket and ship logs out. Simplest to retrofit onto an existing stack with no per-service changes.
  • Native driver pattern: set each container’s logging block directly, e.g. driver: gelf with gelf-address: udp://graylog-host:5140. More explicit and avoids the sidecar entirely, but means touching every compose file instead of adding one new service.

For a homelab with a dozen-plus containers already running, the sidecar pattern is almost always less work - one new service picks up everything already running, instead of editing every existing one.

Retention: the setting that actually needs a decision

Unlike metrics, where a few weeks of history is plenty, logs grow fast and most of them you’ll never read. Decide retention deliberately instead of letting the default run until the disk fills:

  • Loki retention is set in its config (limits_config.retention_period) or via a compactor with retention enabled - without one of those set, Loki keeps everything forever by default.
  • Graylog/OpenSearch retention is managed through Index Set rotation and retention settings in the web UI - rotate daily or weekly, keep a fixed number of indices, and older ones get deleted automatically.

Seven to fourteen days is plenty for a homelab. You’re using this to debug something that just happened, not building a compliance audit trail. If a specific incident needs longer-term evidence, export that one log window before it rotates out rather than keeping everything indefinitely by default.

What to actually run

If Grafana’s already part of the stack, Loki is close to free to add and keeps everything in one pane of glass with the metrics dashboards already there. If logging feels like it deserves its own dedicated tool with a purpose-built search UI and you’ve got the RAM to spare, Graylog is the more complete answer without the full weight of hand-rolled Elasticsearch and Logstash. Skip ELK/Elastic entirely unless there’s a specific reason (usually: matching a work stack) to take on the heaviest option for a job a homelab’s log volume doesn’t actually require.

Either way, the real win isn’t the dashboard, it’s the next time something breaks and the fix is a thirty-second search instead of a guessing tour through docker logs on every container that might be the culprit.