Published Oct 2, 2026, 7:30 AM EDT Shekhar Vaidya is a veteran technology journalist and computer science engineer. He is the founder of TechLatest, where he has spent years providing technical analysis on hardware and Windows ecosystems. Now a Computing Writer at XDA, Shekhar leverages his deep background in NAS, storage solutions, and PC internals to help readers master their tech. One morning, you wake up, open a service, only to find it inaccessible. And when you check your Docker containers or monitoring services, everything looks fine. I have faced it many times in the past, and most recently a few weeks ago with my Nextcloud service. A tool auto-updated the Nextcloud container, but the database needed a manual migration. So, none of my monitoring services detected anything because the Docker container was up and running. That was when I decided to fix this by adding my own health check. Why "running" doesn't mean "working" Everything said "Up," but nothing worked. I have been into homelabbing long enough to understand that if Docker says my container is up and running, that isn’t the final check. Docker doesn’t always know what’s happening inside the container; it only tracks the container’s lifecycle, whether the main process is alive. I am not saying this from some theory I created myself or read somewhere. Similar things have happened to me more than once. I don’t exactly remember all the instances, but I can tell you about a recent one. A few weeks ago, I was testing Watchtower. It was a service to automate Docker container image updates. On paper, it looked good; it removed the boring task of updating each container. I deployed it and let it run for a few days. Everything was good; it was updating all the containers. But one morning I opened my Nextcloud service in the browser and found a database error. What happened was Watchtower pulled a new image, and that image’s code was ahead of the database schema, which threw a schema-breaking version error. The fix was easy: I ran docker exec nextcloud php occ upgrade, and it was done. But the fix isn’t the story. There was an issue with the service, but since Docker showed it up and running, my monitoring service didn’t catch it. The point is the container was up, but the application was broken. I recreated the failure with a two-minute test. I started by creating a throwaway Nginx container and breaking the app from inside, but not the container. The curl command showed it as broken, but docker ps showed it as up and running. This is something you can test yourself. docker run -d --name up-but-dead -p 18081:80 nginx curl -I localhost:18081 docker exec up-but-dead sh -c 'rm /etc/nginx/conf.d/default.conf && nginx -s reload' curl -I localhost:18081; docker ps --filter name=up-but-dead So, should Docker’s default monitoring be enough? In my case, I added another layer to check the health of the service, not just the container. What my containers were already telling me Ten of my containers were already checking themselves Obviously, without finding the current health situation, adding a health check to all my containers wasn’t the right move. So, I started by checking which containers already had a built-in health check. Two simple docker ps commands with health grep showed me the whole picture. Ten out of my twenty-eight running containers already had a health check. They had already been reporting health information for weeks. The remaining eighteen containers didn’t have anything, and as I suspected, Nextcloud was one of them. After that, I ran two simple loop scripts with the docker inspect and docker exec commands. The first result gave the check command, interval, and retries for each one. Basically Docker's built-in checks. A few of my major findings for the ten containers were that most of the containers were configured to surface a failure after 90 seconds (30s × 3 tries), and Immich made it more difficult; it would take around 15 minutes to actually report a failure. For the remaining eighteen containers, I checked for curl or wget inside each. Most of them had at least one; a few, including nextcloud_db, had neither. Portainer and beszel-agent don't even have a shell available. After these findings, I only had to add a health check to Docker compose files and redeploy them. For example, in the Stirling PDF compose file, I added these lines. labels: - autoheal=true healthcheck: test: ["CMD-SHELL", "curl -fsS http://localhost:8080/api/v1/info/status || exit 1"] interval: 30s timeout: 5s retries: 3 start_period: 90s At this point, most of my containers were giving me health information. Ten of my containers already had a built-in health check, and for the others I added my own, but what happens after that? I still needed something that acted on it. Letting the homelab fix itself I broke a service on purpose and then it got interesting This is much more interesting than the discovery and the health check. I added a watcher to the whole setup. It's a small script that runs nonstop under systemd. Before you question me, let me be clear: it doesn’t diagnose a bad config and fix it automatically. The script listens to Docker's health-status event stream. When the status is unhealthy, it restarts the container. It stores each restart in its memory, and maximum restarts are limited to three within 30 minutes to guard the loop. It reacts to the health check events and only acts if autoheal=true is defined on the container. The autoheal label safeguards major containers and databases because without it the script won’t act on them. Each restart, and the moment the loop guard gives up, is notified on my ntfy channel. But to test it out, I actually had to break one of the containers. I froze Stirling PDF because a crash is already handled by Docker's restart policy. And freezing proves my earlier observation that the process is alive, but the application isn't responding properly. As soon as I froze the Java process inside the container, in around 90 seconds the container became unhealthy, and the watcher restarted it in another 10 seconds. And in another 30 seconds, the container became healthy again. So, it took around two minutes to recover the service. But as mentioned before, it doesn’t diagnose and fix a bad config. For example, the Nextcloud case. If the script had been live at that time, it could only have restarted the container and then left it as it is, since a restart can’t run the migration that the Nextcloud fix needed. Self-restarting isn't the same thing as self-healing. Maybe, in the future, I can upgrade the script by adding a small local LLM, but that is a project for later. For now, it does the easy part and leaves the tough part for me to diagnose and fix. I know it doesn’t automate the whole process, but something is better than nothing. Half-baked, but it's staying The watcher is worth keeping in my homelab, but it is currently half-baked. It does the boring part of checking the labeled container’s health and restarts the container if it turns unhealthy. But the improvement possibilities are many, like the local LLM idea I mentioned, and even more. The script part is optional, but a proper health check is definitely a must-have if you self-host several services.
I added health checks to every Docker container, and now my homelab restarts its own broken services
Full Article
Original Source
Read the full article at Xda-developers →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.