PM2 Logs Monitoring and Health Checks in Docker Compose

Tech dashboard illustrating Node.js, PM2-runtime, Docker containers, streaming stdout logs, health checks, and observability graphs in a neon blue-green palette.

If your Node.js services run under PM2 inside Docker Compose, you already value stability and zero-downtime restarts. But reliability doesn’t happen by accident. Clear logs, actionable monitoring, and trustworthy health checks are the guardrails that keep production calm. In this guide, you will learn pragmatic ways to wire PM2 with Docker Compose so your services are observable, self-healing, and ready for scale.

Table of Contents

Why Logs, Monitoring, and Health Checks Matter

Logs tell you what happened. Monitoring tells you what is happening now. Health checks tell Docker when to step in and restart something that can’t heal itself. Together, they reduce mean time to detection (MTTD) and mean time to recovery (MTTR), directly influencing uptime and customer trust.

Inside containers, the rules change slightly: you want processes to write to stdout/stderr, you lean on the platform for rotation and shipping, and you codify liveness/readiness. PM2 adds process supervision and clustering on top, but you need to configure it the “container-native” way.

How PM2 Fits Inside Containers

PM2 can manage one or many Node.js processes. In containers, the recommended entrypoint is pm2-runtime, which behaves well as PID 1 and relays logs to the container’s stdout/stderr.

pm2 vs pm2-runtime

  • pm2: great for VMs or dev machines; requires a daemon and may mask exit codes inside containers.
  • pm2-runtime: designed for Docker/Kubernetes; forwards logs, respects signals, and exits when processes die.

Graceful shutdown

Ensure your app listens to SIGINT/SIGTERM and closes servers and DB connections. PM2 will forward signals so Docker can stop quickly during deploys.

Clustering

For CPU-bound workloads or parallel I/O, PM2’s cluster mode (instances: “max” or a number) balances workers behind a single port. Make sure readiness checks reflect readiness of the whole service, not just one worker.

Logging Strategies for PM2 in Docker Compose

1) Prefer stdout/stderr over files

In containers, write logs to stdout/stderr. PM2 does this by default with pm2-runtime. This lets Docker’s logging driver collect, rotate, and ship logs without extra agents in the container.

  • Use structured logs (JSON) to make searching easier.
  • Include a minimal set of fields: timestamp, level, message, requestId, service, version.

2) Format and enrich logs with environment

  • Set environment variables like NODE_ENV, SERVICE_NAME, and VERSION to add context to every log line.
  • In PM2 ecosystem files, you can enable merge_logs and set log_date_format if you still write files. For containers, prefer JSON logs via your logger (e.g., pino, Winston).

3) Control size with Docker log options

If you rely on the default json-file driver, configure rotation in Compose so logs don’t fill disks:

  • logging.driver: “json-file”
  • logging.options.max-size: “10m”
  • logging.options.max-file: “3”

This rotates container logs automatically without changing PM2.

4) PM2 Logrotate (when you write files)

If you intentionally write to files (e.g., persistent audit logs), use pm2 install pm2-logrotate and configure rotation limits. Keep in mind that file-based logs inside containers are an anti-pattern unless mounted to a volume and shipped externally.

5) Persist logs sparingly

  • Mount a volume if compliance requires retention. Common PM2 paths include ~/.pm2/logs.
  • Ship logs to an external store (ELK, Loki, or cloud) as the source of truth.

Docker Compose Examples and Patterns

Service with pm2-runtime and JSON logs

Key lines you might put under a service in docker-compose.yml:

  • command: [“pm2-runtime”, “start”, “ecosystem.config.js”, “–env”, “production”]
  • environment:
  • – NODE_ENV=production
  • – SERVICE_NAME=api
  • – VERSION=1.3.0
  • logging:
  • driver: “json-file”
  • options:
  • max-size: “10m”
  • max-file: “3”

PM2 ecosystem essentials

In ecosystem.config.js, consider:

  • name: identify the process across environments.
  • script: your entry file (e.g., server.js).
  • instances: e.g., “max” for cluster across cores.
  • exec_mode: “cluster” or “fork”.
  • env_production: environment vars for prod.
  • Avoid out_file/error_file for containers; let pm2-runtime stream logs to stdout/stderr.

Healthcheck with HTTP

Compose snippet lines for an HTTP health endpoint at /health:

  • healthcheck:
  • test: [“CMD”, “curl”, “-f”, “http://localhost:3000/health”]
  • interval: 15s
  • timeout: 3s
  • retries: 5
  • start_period: 20s

Centralized Logging Options

Production systems benefit from external log aggregation for search, dashboards, and alerts.

Use Docker logging drivers

  • gelf: ship to Graylog or Logstash GELF input.
  • syslog: send to a remote syslog.
  • fluentd: forward to Fluentd/Fluent Bit stacks.
  • loki: push logs to Grafana Loki via Promtail or a Loki driver.

Example lines for a GELF driver:

  • logging:
  • driver: gelf
  • options:
  • gelf-address: udp://graylog:12201
  • tag: {{.ImageName}}/{{.Name}}

Don’t duplicate

Pick one path to aggregation to avoid paying twice in egress and storage. Either ship from Docker or from a sidecar agent, not both.

Monitoring Options for PM2 Apps

PM2 built-ins

  • pm2 status and pm2 monit show memory/CPU and restarts. From outside, run docker compose exec service pm2 status.
  • PM2 I/O (pm2.io) offers dashboards, alerts, and tracing. See pm2.keymetrics.io.

Metrics-first monitoring

  • Expose application metrics (e.g., Prometheus at /metrics) for latency, throughput, and errors.
  • Scrape Docker daemon metrics and container stats for CPU/memory saturation insights.
  • Graph with Grafana and set SLO-based alerts.

What to alert on

  • Error rate spikes (5xx, unhandled exceptions).
  • High restart counts or crash loops in PM2.
  • Memory growth over time (possible leak).
  • Health check failures or high tail latency.

Health Checks That Actually Work

Docker Compose health checks determine if a container is “healthy.” They do not restart a container by themselves, but unhealthy containers can trigger retries in orchestrators or be used to gate dependencies.

Liveness vs readiness

  • Liveness: is the process alive? If not, restart it.
  • Readiness: is the app ready to serve? Delay traffic until ready.

In Compose, there’s a single healthcheck. Implement your endpoint so it returns non-200 until the app is truly ready (DB connected, caches warm).

HTTP health

  • Endpoint: /health or /_health.
  • Return 200 and a minimal JSON body like {“status”:”ok”}.
  • Return non-2xx if dependencies are down.

TCP or command health

  • Use test: [“CMD”, “nc”, “-z”, “localhost”, “3000”] to check port availability.
  • Or test: [“CMD-SHELL”, “node healthcheck.js”] for custom logic.

Tuning

  • interval: how often to probe (e.g., 15s).
  • timeout: when a probe is considered failed (e.g., 3s).
  • retries: resilience to transient blips (e.g., 5).
  • start_period: grace during cold start (e.g., 20s+ if migrations run).

Restarts, Dependencies, and Startup Ordering

Restart policies

  • restart: on-failure for batch jobs or scripts.
  • restart: unless-stopped for services you want always-on.

PM2 adds an internal restart policy, but your container still needs a Compose-level policy to recover from process exit or OOMKills.

depends_on with health

Compose v3+ supports dependency conditions. Example lines:

  • depends_on:
  • db:
  • condition: service_healthy

This blocks the app until the database is healthy. Still keep retry logic in your app; don’t rely solely on ordering.

Quick Observability Playbook

1) Baseline logging

  • Run with pm2-runtime and write structured logs to stdout.
  • Enable Docker log rotation (max-size, max-file).
  • Tag logs with SERVICE_NAME and VERSION.

2) Metrics and dashboards

  • Instrument latency, throughput, and error rate.
  • Expose /metrics and scrape with Prometheus.
  • Dashboard with Grafana; set SLOs and alerts.

3) Health and restarts

  • Create a robust /health endpoint.
  • Add a Compose healthcheck with start_period tuned to boot time.
  • Use restart: unless-stopped for services.

4) Centralize logs

  • Pick a logging driver (e.g., gelf, fluentd) or a sidecar.
  • Ship logs off-host; keep disk usage predictable.

Troubleshooting Common Issues

Logs are missing or duplicated

  • Ensure you use pm2-runtime, not pm2, in containers.
  • Remove out_file/error_file from the PM2 config to avoid double logging.
  • Verify only one aggregation path (Docker driver or sidecar), not both.

Health check keeps flapping

  • Increase start_period to cover warm-up time.
  • Probe less often or extend timeout if under heavy load.
  • Ensure the endpoint checks dependencies realistically but caches results briefly to avoid self-DOS.

High CPU/memory but no obvious errors

  • Use pm2 monit to spot workers with leaks.
  • Enable heap snapshots or CPU profiles on demand.
  • Check for unbounded logs or chatty debug modes in production.

Crashes during deploys

  • Implement graceful shutdown; close servers on SIGTERM.
  • Delay traffic with readiness checks until warm.
  • Use max_restarts in PM2 if a script rapidly fails.

Security and Performance Considerations

Protect sensitive data

  • Redact secrets before logging. Never log access tokens or passwords.
  • Limit request/response bodies; log IDs, not payloads.

Keep logs affordable

  • Set sampling for high-volume routes.
  • Expire debug logs quickly; keep only what you need for audits and forensics.

Resource limits

  • In Compose, use deploy.resources.limits or Docker run flags to constrain CPU/memory.
  • Alert on OOMKills; verify memory ceilings for Node.js (–max-old-space-size as needed).

Conclusion

PM2 and Docker Compose can deliver a robust platform for Node.js services—if you treat logs, monitoring, and health checks as first-class citizens. Stream logs to stdout, rotate and centralize them, instrument metrics that map to user experience, and codify readiness with honest health endpoints. With these pieces in place, your apps recover faster, alert earlier, and scale with confidence.

Frequently Asked Questions

Should I use pm2 or pm2-runtime in Docker?

Use pm2-runtime. It behaves as PID 1, forwards logs to stdout/stderr, and exits correctly so Docker can apply restart policies.

Where do PM2 logs go inside a container?

With pm2-runtime, logs go to stdout/stderr. View them via docker compose logs -f service. Only write files if you must, and then rotate aggressively or mount a volume.

PM2 logrotate vs Docker log rotation—do I need both?

Use Docker rotation for stdout/stderr. Use PM2 logrotate only if you deliberately write file logs from PM2. Avoid double rotation for the same stream.

How do I make a reliable health check?

Expose a lightweight /health that returns non-200 until dependencies are ready. In Compose, configure healthcheck with sensible start_period, interval, and retries so brief spikes don’t mark the service unhealthy.

Leave a Reply

Your email address will not be published. Required fields are marked *