High-density cPanel & WHM hosting servers host dozens or hundreds of mission-critical customer websites, e-commerce stores, and mailboxes. When a core service like MySQL, PHP-FPM, Apache, or Exim hangs or exhausts available worker sockets, immediate automated remediation is vital to maintain 99.99% service availability.
In the cPanel ecosystem, service self-healing is governed by chkservd (the cPanel Check Services Daemon). chkservd polls configured daemon ports and local sockets on a recurring loop. If a service fails to respond within its timeout threshold, chkservd automatically attempts to restart the process and dispatches alert notifications to system administrators.
However, default configurations often cause severe issues:
- Restart Loops & Flapping: A server experiencing temporary I/O wait can be pushed into a catastrophic crash cycle if
chkservdrepeatedly restarts an active MariaDB process performing crash recovery. - Missing Custom Services: Modern stacks frequently run non-standard daemons—such as Redis, Meilisearch, Memcached, or Node.js background workers—that default cPanel installations completely ignore.
In this operational guide, we will walk through tuning chkservd heartbeats, writing custom service probes in /etc/chkserv.d/, and stabilizing auto-healing on enterprise bare-metal Dedicated Servers and Dedicated Servers in Pakistan.
1. How chkservd Works Under the Hood
The chkservd engine is managed by systemd as cpanel-chkservd.service. Every polling interval (defaulting to 300 seconds, or 5 minutes), chkservd reads its configuration from:
/etc/chkserv.d/chkservd.conf: Master list of enabled/disabled service checks./etc/chkserv.d/{service_name}: Individual service definitions specifying port probes, challenge-response strings, and process matching patterns.
+--------------------------------------------------------------+
| Systemd (PID 1) |
+------------------------------+-------------------------------+
|
v
+--------------------------------------------------------------+
| cpanel-chkservd.service (Loop) |
| - Polls TCP / Unix Sockets defined in /etc/chkserv.d/ |
| - Validates Protocol Handshakes (SMTP, HTTP, POP3, IMAP) |
| - Triggers /scripts/restartsrv if timeouts exceed limits |
+--------------------------------------------------------------+
Checking chkservd Real-Time Status
Run the following commands via SSH as root to inspect the live status of the service monitor:
# Check systemd status
systemctl status cpanel-chkservd
# Run an immediate manual check of all services
/scripts/restartsrv_chkservd --status
# Tail the real-time service check log
tail -f /var/log/chkservd.log
A healthy log excerpt appears as:
Service check ....
[+] cpdavd (127.0.0.1:2077) ... ok
[+] exim (127.0.0.1:25) ... ok
[+] httpd (127.0.0.1:80) ... ok
[+] mysql (127.0.0.1:3306) ... ok
[+] named (127.0.0.1:53) ... ok
2. Adding a Custom Service Monitor (Example: Redis)
Suppose you are running a dedicated Redis instance on port 6379 for WordPress object caching. If Redis crashes, WordPress sites immediately throw 500 errors. We can configure chkservd to monitor Redis and auto-restart it if it ever fails.
Step 1: Create the Service Definition
Create a new file /etc/chkserv.d/redis:
nano /etc/chkserv.d/redis
Add the following protocol definition string:
service[redis]=6379,PING,+PONG,/usr/bin/systemctl restart redis,redis,root
Let’s dissect this directive format:
6379: The TCP portchkservdwill connect to (127.0.0.1:6379).PING: The challenge payload sent down the socket.+PONG: The expected string in the server’s reply./usr/bin/systemctl restart redis: The command to execute if the check fails.redis: The process name inps auxto match.root: The user under which the restart command must run.
Step 2: Enable the Service in chkservd.conf
Append the service flag to /etc/chkserv.d/chkservd.conf:
echo "redis:1" >> /etc/chkserv.d/chkservd.conf
Step 3: Register and Restart chkservd
Notify cPanel of the new service and restart the monitoring daemon:
/scripts/restartsrv_chkservd
Verify in /var/log/chkservd.log that redis is now actively polled:
grep -i "redis" /var/log/chkservd.log
# Expected: [+] redis (127.0.0.1:6379) ... ok
3. Tuning Heartbeat Intervals & Preventing Restart Storms
On high-concurrency shared servers, a spike in database lock contention or CPU usage during backup routines can cause MySQL or Apache to take longer than usual to respond. If chkservd is too aggressive, it will terminate a healthy, processing daemon—corrupting transactional memory tables.
Tuning the Heartbeat Frequency
By default, WHM sets the polling loop to 300 seconds. To adjust this:
- Log in to WHM » Server Configuration » Tweak Settings.
- Navigate to the System tab.
- Locate The number of times chkservd attempts to restart a service before disabling it and Service check timeout.
Alternatively, manage this programmatically via /etc/cpupdate.conf and /var/cpanel/sysinfo.config:
# Inspect current service timeout thresholds (in seconds)
whmapi1 get_tweaksetting key=chkservd_timeout
Creating Custom Restart Wrapper Scripts with Cooldown Throttling
To prevent chkservd from restarting a service during temporary high I/O:
#!/usr/bin/env bash
# /usr/local/sbin/safe_mysql_restart.sh
# Check if system load average is above CPU core count
LOAD=$(awk '{print $1}' /proc/loadavg | cut -d. -f1)
CORES=$(nproc)
if [ "$LOAD" -gt "$((CORES * 2))" ]; then
echo "Load ($LOAD) exceeds safe threshold ($((CORES * 2))). Delaying MySQL restart to prevent disk thrashing." >> /var/log/safe_restart.log
exit 0
fi
# Execute safe graceful restart
systemctl restart mariadb
Point /etc/chkserv.d/mysql to this script to guarantee safe execution.
4. Best Practices for Production Hosting Fleets
| Configuration Area | Recommended Practice | Failure Mode if Ignored |
|---|---|---|
| Database Checks | Use gentle ping challenge (3306) |
Corrupted InnoDB ibdata files during forced restart |
| Web Server | Pair with PHP-FPM Emergency Restart | Cascading 502 Bad Gateway outages |
| DNS Subsystem | Ensure synchronization with cPanel DNS Clusters | Stale DNS zone propagation |
| Notification Flood | Configure notification thresholds in WHM Contact Manager | Admin email inboxes flooded with transient alerts |
Eliminate Service Downtime on Dedicated Bare-Metal Servers
Protect your high-traffic cPanel hosting infrastructure from resource starvation. Nextgen's high-frequency AMD EPYC and Intel Xeon dedicated servers feature enterprise NVMe Gen4 arrays, unthrottled CPU cores, and 24/7 proactive hardware monitoring in Karachi and Islamabad.
