Docker Compose Healthcheck Failing? Diagnose and Fix
When a Docker Compose service is marked unhealthy but application logs show nothing, the healthcheck itself is the problem. This guide shows how to extract the healthcheck definition, run it manually, and fix the failure on Linux and Windows.
Rao Aadil, India 9 min read
Confirm the unhealthy status and where Docker reports it
When a Docker Compose healthcheck failing leaves a service marked unhealthy, the first step is to confirm what Docker itself reports. Run docker ps and look at the STATUS column. On both Linux and Windows, an unhealthy container shows something like Up 2 minutes (unhealthy).
The container status tells you that the healthcheck has failed, but not why. To see the healthcheck state in detail, run:
docker inspect --format '{{json .State.Health}}' <container>
On Linux, macOS, and PowerShell, single quotes around the format string work as written. In Command Prompt, use double quotes:
docker inspect --format "{{json .State.Health}}" <container>
The output includes fields like Status (starting, healthy, or unhealthy), FailingStreak, and Log, which is an array of recent healthcheck executions. Each log entry has the exit code and the output from that run.
If you are working in a Compose project, docker compose ps shows health status for all services. On Linux:
cd /opt/myapp
docker compose ps
On Windows (PowerShell or cmd):
cd C:\opt\myapp
docker compose ps
One thing that surprises many operators: the output from a healthcheck does not go to the container's standard output. Docker stores it in the healthcheck log you just saw under .State.Health.Log. That is why the application logs may show no errors even though the healthcheck is failing. The app never sees the failing request because the healthcheck command itself is the problem.
“It works on one machine and not another, and the error mentions a file or a port rather than your code.”
It reads the compose file, resolves the project name the way Compose does, and checks whether the port is taken, the bind-mount path exists, and the image is present.
It reads the real state of the server before it says anything, shows you the exact command, and waits for your approval. Windows and Linux, over the SSH access you already have.
Find the exact healthcheck definition
Next, find what command Docker is actually running. If the service has a healthcheck in the compose file, it is under the service definition in docker-compose.yml or compose.yaml. A typical definition looks like this:
services:
web:
image: nginx
healthcheck:
test: ["CMD-SHELL", "curl -f http://localhost/health || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 40s
The test is the command to run. interval is how often Docker runs it, timeout is how long to wait for each execution, retries is how many consecutive failures mark the container unhealthy, and start_period gives the container time to start before failures count.
If the compose file has no healthcheck, the image may have one built in through its Dockerfile HEALTHCHECK. To inspect the image:
docker inspect --format '{{json .Config.Healthcheck}}' <image>
Again, use double quotes in Command Prompt. The output shows the Test array, Interval, Timeout, Retries, and StartPeriod.
Keep in mind that a healthcheck defined in compose overrides any healthcheck defined in the Dockerfile. If you add a compose healthcheck but forget to include the original test (or use a completely different command), the container may become unhealthy even though the built-in check was fine.
On Windows, there are extra details. Docker healthchecks on Windows containers do not use /bin/sh. You need cmd /c or powershell in the test. Line endings matter: a script referenced by the healthcheck that was created on Linux with LF may need CRLF on Windows, or it may fail to run. Paths are also Windows-style, so a command that references /opt/myapp/healthcheck.sh must become C:\opt\myapp\healthcheck.ps1.
Run the healthcheck manually to capture exit code and output
The healthcheck log in docker inspect gives the exit code and output, but running the command yourself inside the container is often faster and lets you capture more context. Start by extracting the exact command from the test definition.
For the compose example above, the command inside the container is effectively:
curl -f http://localhost/health
To run it manually in a Linux container:
docker exec <container> sh -c "curl -f http://localhost/health"
For a Windows container, use cmd or PowerShell:
docker exec <container> cmd /c "curl -f http://localhost/health"
or
docker exec <container> powershell -Command "Invoke-WebRequest -UseBasicParsing http://localhost/health"
After the command, capture the exit code immediately. In a Linux shell or bash:
echo $?
In PowerShell, after running docker exec, use:
$LASTEXITCODE
In Command Prompt:
echo %ERRORLEVEL%
The exit code is the single most important piece of information. A non-zero exit code means failure; zero means success. The output you see from the manual run will show the actual error: connection refused, 404 Not Found, command not found, a timeout, a DNS failure, or a missing environment variable. This is the same output Docker puts in the healthcheck log, but by running it interactively you can see it in real time and avoid the delay between scheduled checks.
Diagnose common healthcheck failure causes
The following are the patterns you will most often find.
-
Missing tool. The healthcheck uses
curlorwget, but the container image does not include it. This happens on slim Linux images and on Windows containers that only include PowerShell. The manual run showscommand not foundorcurl: not found. -
Wrong endpoint or path. The URL points to a route that does not exist. For example,
curl -f http://localhost/healthreturns404 Not Foundbecause the app only exposes/healthzor/status. The request never reaches a route that logs an error, so application logs stay clean. -
Network not ready. In Compose, a service may depend on a database or another service that is still starting. The healthcheck runs immediately and fails with
connection refused. Setting a generousstart_periodandretriesprevents these transient failures from marking the container unhealthy. -
Environment variables not set. The healthcheck references
$HOSTor$PORT, but those variables are not defined in the container environment. The command fails because the shell expands them to empty strings. -
File path mismatch. On Windows containers, a healthcheck that references a script or binary with a Linux path (
/opt/myapp/check.sh) will fail. Use Windows paths (C:\opt\myapp\check.ps1) and make sure the file is inside the container at that location. -
Script permissions. On Linux, a healthcheck that points to a script may fail because the script is not executable. The healthcheck log will show
permission denied. Fix withchmod +xin the image or usesh /path/to/script. On Windows, a script may need an explicit interpreter likepowershell -File C:\script.ps1. -
Exit code semantics. Some commands return a non-zero exit code on success because of incorrect flags or shell handling. For example,
curl http://localhost/healthwithout-freturns 0 even on a 404, so the healthcheck passes incorrectly. Always use-f(or equivalent) to make HTTP errors fail the check. -
Healthcheck interval too short. If
intervalis too short andretriesis small, a slow first start or a brief network blip can cause the container to flap between healthy and unhealthy, and eventually settle on unhealthy.
Fix the healthcheck for both Linux and Windows
The fix usually means editing the compose file with a corrected healthcheck and then recreating the container.
For Linux containers, a robust healthcheck using curl is:
healthcheck:
test: ["CMD-SHELL", "curl -f http://localhost/health || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 40s
If the image does not have curl but has wget, use:
test: ["CMD-SHELL", "wget -q --spider http://localhost/health || exit 1"]
For Windows containers, prefer PowerShell. Use the array form of test to avoid shell quoting issues:
healthcheck:
test: ["CMD", "powershell", "-Command", "try { Invoke-WebRequest -UseBasicParsing http://localhost/health } catch { exit 1 }"]
interval: 30s
timeout: 10s
retries: 3
start_period: 40s
If using cmd instead, something like this works:
test: ["CMD", "cmd", "/c", "curl -f http://localhost/health || exit 1"]
Make sure the command is correctly formatted. In a Dockerfile, HEALTHCHECK uses CMD syntax. In a compose file, test can use CMD-SHELL (which wraps the string in /bin/sh -c on Linux or cmd /S /C on Windows) or an array of CMD and arguments. CMD-SHELL is convenient for simple commands, but if you need PowerShell on Windows, use the CMD array with powershell and -Command.
After editing the compose file, apply the change by recreating the container. A healthcheck change alone may not be picked up by a simple restart in all cases; the safest approach is:
docker compose up -d --force-recreate
On both Linux and Windows, this recreates the containers with the new definition. If you only need to restart a single service:
docker compose restart <service>
Then verify the new status with docker ps or:
docker inspect --format '{{json .State.Health}}' <container>
You should eventually see "Status":"healthy".
Prevent future healthcheck failures
The best prevention is to test the healthcheck command manually inside the container before you ever add it to the compose file. That catches missing tools, wrong paths, and permission issues immediately.
Use a generous start_period for slow-starting applications and a reasonable retries count. A healthcheck that is too strict will cause false unhealthy states and paging at 3 a.m.
Keep healthchecks simple and idempotent. A one-line curl -f or Invoke-WebRequest is far easier to debug than a multi-line script.
Document the healthcheck's dependencies: required tools, network endpoints, and environment variables. If another engineer needs to move the service to a different base image, that documentation will prevent a broken healthcheck.
Monitor unhealthy containers across all your servers. If you use an AI agent to manage Docker, it can inspect healthcheck definitions and run the failing healthcheck command inside a container (with your approval) to capture the exit code and output. It cannot edit the healthcheck definition or compose file itself, but after you fix the file, it can restart the stack with the correct project name.