{"id":873,"date":"2026-07-16T20:47:23","date_gmt":"2026-07-16T20:47:23","guid":{"rendered":"https:\/\/blog.asambe.ai\/index.php\/2026\/07\/16\/pm2-logs-monitoring-and-health-checks-in-docker-compose\/"},"modified":"2026-07-16T20:47:26","modified_gmt":"2026-07-16T20:47:26","slug":"pm2-logs-monitoring-and-health-checks-in-docker-compose","status":"publish","type":"post","link":"https:\/\/blog.asambe.ai\/index.php\/2026\/07\/16\/pm2-logs-monitoring-and-health-checks-in-docker-compose\/","title":{"rendered":"PM2 Logs Monitoring and Health Checks in Docker Compose"},"content":{"rendered":"<p>If your Node.js services run under PM2 inside Docker Compose, you already value stability and zero-downtime restarts. But reliability doesn\u2019t happen by accident. Clear logs, actionable monitoring, and trustworthy health checks are the guardrails that keep production calm. In this guide, you will learn pragmatic ways to wire PM2 with Docker Compose so your services are observable, self-healing, and ready for scale.<\/p>\n<h2>Table of Contents<\/h2>\n<ul>\n<li><a href=\"#why-it-matters\">Why Logs, Monitoring, and Health Checks Matter<\/a><\/li>\n<li><a href=\"#pm2-in-containers\">How PM2 Fits Inside Containers<\/a><\/li>\n<li><a href=\"#logging-strategies\">Logging Strategies for PM2 in Docker Compose<\/a><\/li>\n<li><a href=\"#compose-examples\">Docker Compose Examples and Patterns<\/a><\/li>\n<li><a href=\"#centralized-logging\">Centralized Logging Options<\/a><\/li>\n<li><a href=\"#monitoring-options\">Monitoring Options for PM2 Apps<\/a><\/li>\n<li><a href=\"#health-checks\">Health Checks That Actually Work<\/a><\/li>\n<li><a href=\"#restarts-dependencies\">Restarts, Dependencies, and Startup Ordering<\/a><\/li>\n<li><a href=\"#observability-playbook\">Quick Observability Playbook<\/a><\/li>\n<li><a href=\"#troubleshooting\">Troubleshooting Common Issues<\/a><\/li>\n<li><a href=\"#security-performance\">Security and Performance Considerations<\/a><\/li>\n<li><a href=\"#conclusion\">Conclusion<\/a><\/li>\n<li><a href=\"#frequently-asked-questions\">Frequently Asked Questions<\/a><\/li>\n<\/ul>\n<h2 id=\"why-it-matters\">Why Logs, Monitoring, and Health Checks Matter<\/h2>\n<p>Logs tell you what happened. Monitoring tells you what is happening now. Health checks tell Docker when to step in and restart something that can\u2019t heal itself. Together, they reduce mean time to detection (MTTD) and mean time to recovery (MTTR), directly influencing uptime and customer trust.<\/p>\n<p>Inside containers, the rules change slightly: you want processes to write to <em>stdout\/stderr<\/em>, you lean on the platform for rotation and shipping, and you codify liveness\/readiness. PM2 adds process supervision and clustering on top, but you need to configure it the \u201ccontainer-native\u201d way.<\/p>\n<h2 id=\"pm2-in-containers\">How PM2 Fits Inside Containers<\/h2>\n<p>PM2 can manage one or many Node.js processes. In containers, the recommended entrypoint is <strong>pm2-runtime<\/strong>, which behaves well as PID 1 and relays logs to the container\u2019s stdout\/stderr.<\/p>\n<h3>pm2 vs pm2-runtime<\/h3>\n<ul>\n<li><strong>pm2<\/strong>: great for VMs or dev machines; requires a daemon and may mask exit codes inside containers.<\/li>\n<li><strong>pm2-runtime<\/strong>: designed for Docker\/Kubernetes; forwards logs, respects signals, and exits when processes die.<\/li>\n<\/ul>\n<h3>Graceful shutdown<\/h3>\n<p>Ensure your app listens to SIGINT\/SIGTERM and closes servers and DB connections. PM2 will forward signals so Docker can stop quickly during deploys.<\/p>\n<h3>Clustering<\/h3>\n<p>For CPU-bound workloads or parallel I\/O, PM2\u2019s cluster mode (<em>instances: &#8220;max&#8221;<\/em> or a number) balances workers behind a single port. Make sure readiness checks reflect <em>readiness<\/em> of the whole service, not just one worker.<\/p>\n<h2 id=\"logging-strategies\">Logging Strategies for PM2 in Docker Compose<\/h2>\n<h3>1) Prefer stdout\/stderr over files<\/h3>\n<p>In containers, write logs to stdout\/stderr. PM2 does this by default with <em>pm2-runtime<\/em>. This lets Docker\u2019s logging driver collect, rotate, and ship logs without extra agents in the container.<\/p>\n<ul>\n<li>Use structured logs (JSON) to make searching easier.<\/li>\n<li>Include a minimal set of fields: timestamp, level, message, requestId, service, version.<\/li>\n<\/ul>\n<h3>2) Format and enrich logs with environment<\/h3>\n<ul>\n<li>Set environment variables like <em>NODE_ENV<\/em>, <em>SERVICE_NAME<\/em>, and <em>VERSION<\/em> to add context to every log line.<\/li>\n<li>In PM2 ecosystem files, you can enable <em>merge_logs<\/em> and set <em>log_date_format<\/em> if you still write files. For containers, prefer JSON logs via your logger (e.g., pino, Winston).<\/li>\n<\/ul>\n<h3>3) Control size with Docker log options<\/h3>\n<p>If you rely on the default <em>json-file<\/em> driver, configure rotation in Compose so logs don\u2019t fill disks:<\/p>\n<ul>\n<li><em>logging.driver: &#8220;json-file&#8221;<\/em><\/li>\n<li><em>logging.options.max-size: &#8220;10m&#8221;<\/em><\/li>\n<li><em>logging.options.max-file: &#8220;3&#8221;<\/em><\/li>\n<\/ul>\n<p>This rotates container logs automatically without changing PM2.<\/p>\n<h3>4) PM2 Logrotate (when you write files)<\/h3>\n<p>If you intentionally write to files (e.g., persistent audit logs), use <em>pm2 install pm2-logrotate<\/em> and configure rotation limits. Keep in mind that file-based logs inside containers are an anti-pattern unless mounted to a volume and shipped externally.<\/p>\n<h3>5) Persist logs sparingly<\/h3>\n<ul>\n<li>Mount a volume if compliance requires retention. Common PM2 paths include <em>~\/.pm2\/logs<\/em>.<\/li>\n<li>Ship logs to an external store (ELK, Loki, or cloud) as the source of truth.<\/li>\n<\/ul>\n<h2 id=\"compose-examples\">Docker Compose Examples and Patterns<\/h2>\n<h3>Service with pm2-runtime and JSON logs<\/h3>\n<p>Key lines you might put under a service in docker-compose.yml:<\/p>\n<ul>\n<li><em>command: [&#8220;pm2-runtime&#8221;, &#8220;start&#8221;, &#8220;ecosystem.config.js&#8221;, &#8220;&#8211;env&#8221;, &#8220;production&#8221;]<\/em><\/li>\n<li><em>environment:<\/em><\/li>\n<li><em>&#8211; NODE_ENV=production<\/em><\/li>\n<li><em>&#8211; SERVICE_NAME=api<\/em><\/li>\n<li><em>&#8211; VERSION=1.3.0<\/em><\/li>\n<li><em>logging:<\/em><\/li>\n<li><em>  driver: &#8220;json-file&#8221;<\/em><\/li>\n<li><em>  options:<\/em><\/li>\n<li><em>    max-size: &#8220;10m&#8221;<\/em><\/li>\n<li><em>    max-file: &#8220;3&#8221;<\/em><\/li>\n<\/ul>\n<h3>PM2 ecosystem essentials<\/h3>\n<p>In <em>ecosystem.config.js<\/em>, consider:<\/p>\n<ul>\n<li><em>name<\/em>: identify the process across environments.<\/li>\n<li><em>script<\/em>: your entry file (e.g., server.js).<\/li>\n<li><em>instances<\/em>: e.g., &#8220;max&#8221; for cluster across cores.<\/li>\n<li><em>exec_mode<\/em>: &#8220;cluster&#8221; or &#8220;fork&#8221;.<\/li>\n<li><em>env_production<\/em>: environment vars for prod.<\/li>\n<li>Avoid <em>out_file<\/em>\/<em>error_file<\/em> for containers; let pm2-runtime stream logs to stdout\/stderr.<\/li>\n<\/ul>\n<h3>Healthcheck with HTTP<\/h3>\n<p>Compose snippet lines for an HTTP health endpoint at <em>\/health<\/em>:<\/p>\n<ul>\n<li><em>healthcheck:<\/em><\/li>\n<li><em>  test: [&#8220;CMD&#8221;, &#8220;curl&#8221;, &#8220;-f&#8221;, &#8220;http:\/\/localhost:3000\/health&#8221;]<\/em><\/li>\n<li><em>  interval: 15s<\/em><\/li>\n<li><em>  timeout: 3s<\/em><\/li>\n<li><em>  retries: 5<\/em><\/li>\n<li><em>  start_period: 20s<\/em><\/li>\n<\/ul>\n<h2 id=\"centralized-logging\">Centralized Logging Options<\/h2>\n<p>Production systems benefit from external log aggregation for search, dashboards, and alerts.<\/p>\n<h3>Use Docker logging drivers<\/h3>\n<ul>\n<li><strong>gelf<\/strong>: ship to Graylog or Logstash GELF input.<\/li>\n<li><strong>syslog<\/strong>: send to a remote syslog.<\/li>\n<li><strong>fluentd<\/strong>: forward to Fluentd\/Fluent Bit stacks.<\/li>\n<li><strong>loki<\/strong>: push logs to Grafana Loki via Promtail or a Loki driver.<\/li>\n<\/ul>\n<p>Example lines for a GELF driver:<\/p>\n<ul>\n<li><em>logging:<\/em><\/li>\n<li><em>  driver: gelf<\/em><\/li>\n<li><em>  options:<\/em><\/li>\n<li><em>    gelf-address: udp:\/\/graylog:12201<\/em><\/li>\n<li><em>    tag: {{.ImageName}}\/{{.Name}}<\/em><\/li>\n<\/ul>\n<h3>Don\u2019t duplicate<\/h3>\n<p>Pick <em>one<\/em> path to aggregation to avoid paying twice in egress and storage. Either ship from Docker or from a sidecar agent, not both.<\/p>\n<h2 id=\"monitoring-options\">Monitoring Options for PM2 Apps<\/h2>\n<h3>PM2 built-ins<\/h3>\n<ul>\n<li><strong>pm2 status<\/strong> and <strong>pm2 monit<\/strong> show memory\/CPU and restarts. From outside, run <em>docker compose exec service pm2 status<\/em>.<\/li>\n<li><strong>PM2 I\/O (pm2.io)<\/strong> offers dashboards, alerts, and tracing. See <a href=\"https:\/\/pm2.keymetrics.io\/\">pm2.keymetrics.io<\/a>.<\/li>\n<\/ul>\n<h3>Metrics-first monitoring<\/h3>\n<ul>\n<li>Expose application metrics (e.g., Prometheus at <em>\/metrics<\/em>) for latency, throughput, and errors.<\/li>\n<li>Scrape Docker daemon metrics and container stats for CPU\/memory saturation insights.<\/li>\n<li>Graph with Grafana and set SLO-based alerts.<\/li>\n<\/ul>\n<h3>What to alert on<\/h3>\n<ul>\n<li>Error rate spikes (5xx, unhandled exceptions).<\/li>\n<li>High restart counts or crash loops in PM2.<\/li>\n<li>Memory growth over time (possible leak).<\/li>\n<li>Health check failures or high tail latency.<\/li>\n<\/ul>\n<h2 id=\"health-checks\">Health Checks That Actually Work<\/h2>\n<p>Docker Compose health checks determine if a container is \u201chealthy.\u201d They do not restart a container by themselves, but unhealthy containers can trigger retries in orchestrators or be used to gate dependencies.<\/p>\n<h3>Liveness vs readiness<\/h3>\n<ul>\n<li><strong>Liveness<\/strong>: is the process alive? If not, restart it.<\/li>\n<li><strong>Readiness<\/strong>: is the app ready to serve? Delay traffic until ready.<\/li>\n<\/ul>\n<p>In Compose, there\u2019s a single <em>healthcheck<\/em>. Implement your endpoint so it returns non-200 until the app is truly ready (DB connected, caches warm).<\/p>\n<h3>HTTP health<\/h3>\n<ul>\n<li>Endpoint: <em>\/health<\/em> or <em>\/_health<\/em>.<\/li>\n<li>Return 200 and a minimal JSON body like <em>{&#8220;status&#8221;:&#8221;ok&#8221;}<\/em>.<\/li>\n<li>Return non-2xx if dependencies are down.<\/li>\n<\/ul>\n<h3>TCP or command health<\/h3>\n<ul>\n<li>Use <em>test: [&#8220;CMD&#8221;, &#8220;nc&#8221;, &#8220;-z&#8221;, &#8220;localhost&#8221;, &#8220;3000&#8221;]<\/em> to check port availability.<\/li>\n<li>Or <em>test: [&#8220;CMD-SHELL&#8221;, &#8220;node healthcheck.js&#8221;]<\/em> for custom logic.<\/li>\n<\/ul>\n<h3>Tuning<\/h3>\n<ul>\n<li><em>interval<\/em>: how often to probe (e.g., 15s).<\/li>\n<li><em>timeout<\/em>: when a probe is considered failed (e.g., 3s).<\/li>\n<li><em>retries<\/em>: resilience to transient blips (e.g., 5).<\/li>\n<li><em>start_period<\/em>: grace during cold start (e.g., 20s+ if migrations run).<\/li>\n<\/ul>\n<h2 id=\"restarts-dependencies\">Restarts, Dependencies, and Startup Ordering<\/h2>\n<h3>Restart policies<\/h3>\n<ul>\n<li><em>restart: on-failure<\/em> for batch jobs or scripts.<\/li>\n<li><em>restart: unless-stopped<\/em> for services you want always-on.<\/li>\n<\/ul>\n<p>PM2 adds an internal restart policy, but your container still needs a Compose-level policy to recover from process exit or OOMKills.<\/p>\n<h3>depends_on with health<\/h3>\n<p>Compose v3+ supports dependency conditions. Example lines:<\/p>\n<ul>\n<li><em>depends_on:<\/em><\/li>\n<li><em>  db:<\/em><\/li>\n<li><em>    condition: service_healthy<\/em><\/li>\n<\/ul>\n<p>This blocks the app until the database is healthy. Still keep retry logic in your app; don\u2019t rely solely on ordering.<\/p>\n<h2 id=\"observability-playbook\">Quick Observability Playbook<\/h2>\n<h3>1) Baseline logging<\/h3>\n<ul>\n<li>Run with <em>pm2-runtime<\/em> and write structured logs to stdout.<\/li>\n<li>Enable Docker log rotation (<em>max-size<\/em>, <em>max-file<\/em>).<\/li>\n<li>Tag logs with <em>SERVICE_NAME<\/em> and <em>VERSION<\/em>.<\/li>\n<\/ul>\n<h3>2) Metrics and dashboards<\/h3>\n<ul>\n<li>Instrument latency, throughput, and error rate.<\/li>\n<li>Expose <em>\/metrics<\/em> and scrape with Prometheus.<\/li>\n<li>Dashboard with Grafana; set SLOs and alerts.<\/li>\n<\/ul>\n<h3>3) Health and restarts<\/h3>\n<ul>\n<li>Create a robust <em>\/health<\/em> endpoint.<\/li>\n<li>Add a Compose healthcheck with <em>start_period<\/em> tuned to boot time.<\/li>\n<li>Use <em>restart: unless-stopped<\/em> for services.<\/li>\n<\/ul>\n<h3>4) Centralize logs<\/h3>\n<ul>\n<li>Pick a logging driver (e.g., gelf, fluentd) or a sidecar.<\/li>\n<li>Ship logs off-host; keep disk usage predictable.<\/li>\n<\/ul>\n<h2 id=\"troubleshooting\">Troubleshooting Common Issues<\/h2>\n<h3>Logs are missing or duplicated<\/h3>\n<ul>\n<li>Ensure you use <em>pm2-runtime<\/em>, not <em>pm2<\/em>, in containers.<\/li>\n<li>Remove <em>out_file<\/em>\/<em>error_file<\/em> from the PM2 config to avoid double logging.<\/li>\n<li>Verify only one aggregation path (Docker driver <em>or<\/em> sidecar), not both.<\/li>\n<\/ul>\n<h3>Health check keeps flapping<\/h3>\n<ul>\n<li>Increase <em>start_period<\/em> to cover warm-up time.<\/li>\n<li>Probe less often or extend <em>timeout<\/em> if under heavy load.<\/li>\n<li>Ensure the endpoint checks dependencies realistically but caches results briefly to avoid self-DOS.<\/li>\n<\/ul>\n<h3>High CPU\/memory but no obvious errors<\/h3>\n<ul>\n<li>Use <em>pm2 monit<\/em> to spot workers with leaks.<\/li>\n<li>Enable heap snapshots or CPU profiles on demand.<\/li>\n<li>Check for unbounded logs or chatty debug modes in production.<\/li>\n<\/ul>\n<h3>Crashes during deploys<\/h3>\n<ul>\n<li>Implement graceful shutdown; close servers on SIGTERM.<\/li>\n<li>Delay traffic with readiness checks until warm.<\/li>\n<li>Use <em>max_restarts<\/em> in PM2 if a script rapidly fails.<\/li>\n<\/ul>\n<h2 id=\"security-performance\">Security and Performance Considerations<\/h2>\n<h3>Protect sensitive data<\/h3>\n<ul>\n<li>Redact secrets before logging. Never log access tokens or passwords.<\/li>\n<li>Limit request\/response bodies; log IDs, not payloads.<\/li>\n<\/ul>\n<h3>Keep logs affordable<\/h3>\n<ul>\n<li>Set sampling for high-volume routes.<\/li>\n<li>Expire debug logs quickly; keep only what you need for audits and forensics.<\/li>\n<\/ul>\n<h3>Resource limits<\/h3>\n<ul>\n<li>In Compose, use <em>deploy.resources.limits<\/em> or Docker run flags to constrain CPU\/memory.<\/li>\n<li>Alert on OOMKills; verify memory ceilings for Node.js (<em>&#8211;max-old-space-size<\/em> as needed).<\/li>\n<\/ul>\n<h2 id=\"conclusion\">Conclusion<\/h2>\n<p>PM2 and Docker Compose can deliver a robust platform for Node.js services\u2014if you treat logs, monitoring, and health checks as first-class citizens. Stream logs to stdout, rotate and centralize them, instrument metrics that map to user experience, and codify readiness with honest health endpoints. With these pieces in place, your apps recover faster, alert earlier, and scale with confidence.<\/p>\n<h2 id=\"frequently-asked-questions\">Frequently Asked Questions<\/h2>\n<p><strong>Should I use pm2 or pm2-runtime in Docker?<\/strong><\/p>\n<p>Use <em>pm2-runtime<\/em>. It behaves as PID 1, forwards logs to stdout\/stderr, and exits correctly so Docker can apply restart policies.<\/p>\n<p><strong>Where do PM2 logs go inside a container?<\/strong><\/p>\n<p>With <em>pm2-runtime<\/em>, logs go to stdout\/stderr. View them via <em>docker compose logs -f service<\/em>. Only write files if you must, and then rotate aggressively or mount a volume.<\/p>\n<p><strong>PM2 logrotate vs Docker log rotation\u2014do I need both?<\/strong><\/p>\n<p>Use Docker rotation for stdout\/stderr. Use PM2 logrotate only if you deliberately write file logs from PM2. Avoid double rotation for the same stream.<\/p>\n<p><strong>How do I make a reliable health check?<\/strong><\/p>\n<p>Expose a lightweight <em>\/health<\/em> that returns non-200 until dependencies are ready. In Compose, configure <em>healthcheck<\/em> with sensible <em>start_period<\/em>, <em>interval<\/em>, and <em>retries<\/em> so brief spikes don\u2019t mark the service unhealthy.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Master PM2 logs, monitoring, and Docker Compose health checks. Configure log rotation, alerts, and reliable restarts to harden production Node.js apps.<\/p>\n","protected":false},"author":1,"featured_media":872,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[8],"tags":[],"class_list":["post-873","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog-posts"],"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"https:\/\/i0.wp.com\/blog.asambe.ai\/wp-content\/uploads\/2026\/07\/2026-07-16-20-47-12-data.png?fit=1024%2C1024&ssl=1","_links":{"self":[{"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/posts\/873","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/comments?post=873"}],"version-history":[{"count":1,"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/posts\/873\/revisions"}],"predecessor-version":[{"id":874,"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/posts\/873\/revisions\/874"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/media\/872"}],"wp:attachment":[{"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/media?parent=873"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/categories?post=873"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.asambe.ai\/index.php\/wp-json\/wp\/v2\/tags?post=873"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}