Anyone can write a script that checks disk space. The difference between that and a monitor people keep enabled is entirely in the alerting discipline: not alerting on a blip, not alerting twice, and saying when it is over.
The Requirements
| # | Requirement | Why it matters |
|---|---|---|
| 1 | Pluggable checks — add one without touching the core | Checks change weekly; the engine should not |
| 2 | Every alert carries diagnostic context | An alert without evidence just makes someone log in |
| 3 | Flap damping: N consecutive failures | A single blip must not page anyone |
| 4 | Deduplication with a cooldown | Repeating every 5 minutes gets the channel muted |
| 5 | Recovery notifications | Otherwise nobody knows it is over |
| 6 | State survives restarts | Damping and dedup need memory |
| 7 | Dead-man switch | Detects the monitor itself dying |
| 8 | Alert delivery never breaks the run | A failed webhook must not abort the checks |
Step 1: Core Engine
#!/usr/bin/env bash
#
# monitor.sh — run health checks, alert with damping and deduplication.
#
set -Eeuo pipefail
readonly SCRIPT_NAME=${0##*/}
readonly SCRIPT_DIR=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
readonly HOSTNAME_FQDN=$(hostname -f 2>/dev/null || hostname)
readonly STATE_DIR="${MONITOR_STATE_DIR:-/var/lib/monitor}"
readonly FAIL_THRESHOLD="${MONITOR_FAIL_THRESHOLD:-3}" # consecutive failures before alerting
readonly COOLDOWN="${MONITOR_COOLDOWN:-3600}" # seconds between repeat alerts
readonly DRY_RUN="${MONITOR_DRY_RUN:-0}"
mkdir -p "$STATE_DIR"
log() { printf '%s [%-5s] %s\n' "$(date -Iseconds)" "$1" "${*:2}" >&2; }
info() { log INFO "$@"; }
warn() { log WARN "$@"; }
# Sanitise a check name into something safe for a filename
_key() { printf '%s' "${1//[^a-zA-Z0-9_-]/_}"; }
The _key helper matters: check names like disk:/var/lib become filenames, and an unsanitised / creates a directory traversal rather than a state file.
Step 2: State — Damping and Deduplication
# Track consecutive failures; return 0 when the threshold is newly reached
record_failure() {
local name; name=$(_key "$1")
local f="${STATE_DIR}/fail-${name}"
local count
count=$(( $(cat "$f" 2>/dev/null || echo 0) + 1 ))
printf '%s\n' "$count" > "$f"
(( count >= FAIL_THRESHOLD ))
}
# Clear the counter; return 0 if this is a RECOVERY from an alerted state
record_success() {
local name; name=$(_key "$1")
local f="${STATE_DIR}/fail-${name}"
local a="${STATE_DIR}/alerted-${name}"
rm -f "$f"
if [[ -f $a ]]; then
rm -f "$a"
return 0 # was alerting, now healthy → recovery
fi
return 1
}
# Cooldown: return 0 only if we are allowed to alert now
should_alert() {
local name; name=$(_key "$1")
local a="${STATE_DIR}/alerted-${name}"
if [[ -f $a ]]; then
local age=$(( $(date +%s) - $(stat -c %Y "$a" 2>/dev/null || stat -f %m "$a") ))
(( age < COOLDOWN )) && return 1
fi
touch "$a"
return 0
}
Three files per check, and the whole alerting policy falls out of them: fail-* counts consecutive failures, alerted-* exists while an alert is outstanding and its mtime is the cooldown clock.
Step 3: Notification
notify() {
local severity=$1 title=$2 body=$3
if (( DRY_RUN )); then
info "DRY-RUN alert [${severity}] ${title}: ${body}"
return 0
fi
local colour
case $severity in
critical) colour="#dc2626" ;;
warning) colour="#f59e0b" ;;
resolved) colour="#10b981" ;;
*) colour="#6b7280" ;;
esac
if [[ -n ${SLACK_WEBHOOK_URL:-} ]]; then
local payload
payload=$(jq -n \
--arg c "$colour" \
--arg t "[${severity^^}] ${HOSTNAME_FQDN}: ${title}" \
--arg b "$body" \
'{attachments:[{color:$c, title:$t, text:$b, ts:(now|floor)}]}')
curl -fsS --max-time 10 -X POST \
-H 'Content-Type: application/json' \
-d "$payload" "$SLACK_WEBHOOK_URL" >/dev/null \
|| warn "slack delivery failed"
fi
if [[ -n ${ALERT_EMAIL:-} ]] && command -v mail >/dev/null; then
printf '%s\n' "$body" | mail -s "[${severity}] ${HOSTNAME_FQDN}: ${title}" "$ALERT_EMAIL" \
|| warn "email delivery failed"
fi
}
Every delivery path ends in || warn. Under set -e an unguarded curl failure would abort the whole run — so the check that found a genuine problem would exit before reporting anything else.
Step 4: Checks as Plugins
Each check is a function returning 0 for healthy, non-zero for failing, and printing its evidence to stdout.
check_disk() {
local threshold="${DISK_THRESHOLD:-85}"
local failed=0
while read -r pct mount; do
if (( pct >= threshold )); then
local top
top=$(du -xh --max-depth=2 "$mount" 2>/dev/null | sort -rh | head -3 |
awk '{printf "%s (%s); ", $2, $1}')
echo "${mount} at ${pct}% [threshold ${threshold}%] — largest: ${top}"
failed=1
fi
done < <(df -hP -x tmpfs -x devtmpfs -x overlay |
awk 'NR>1 {gsub(/%/,"",$5); print $5, $6}')
return "$failed"
}
check_memory() {
local threshold="${MEM_THRESHOLD:-90}"
local total avail used_pct
read -r total avail < <(free -m | awk '/^Mem:/ {print $2, $7}')
used_pct=$(( (total - avail) * 100 / total ))
(( used_pct < threshold )) && return 0
local top
top=$(ps -eo comm,rss --sort=-rss | awk 'NR>1 && NR<=4 {printf "%s(%dMB) ", $1, $2/1024}')
echo "memory ${used_pct}% used [threshold ${threshold}%] — top: ${top}"
return 1
}
check_service() {
local svc=$1
systemctl is-active --quiet "$svc" && return 0
local state logs
state=$(systemctl is-active "$svc" 2>&1 || true)
logs=$(journalctl -u "$svc" -n 3 --no-pager -o cat 2>/dev/null | tr '\n' '; ')
echo "service ${svc} is ${state} — recent: ${logs}"
return 1
}
check_http() {
local url=$1 expect="${2:-200}"
local code time_total
read -r code time_total < <(
curl -sS -o /dev/null -w '%{http_code} %{time_total}' \
--max-time 10 "$url" 2>/dev/null || echo "000 0"
)
[[ $code == "$expect" ]] && return 0
echo "${url} returned ${code} (expected ${expect}) after ${time_total}s"
return 1
}
check_cert_expiry() {
local host=$1 days="${CERT_WARN_DAYS:-14}"
local end_date end_epoch remaining
end_date=$(echo | openssl s_client -servername "$host" -connect "${host}:443" 2>/dev/null |
openssl x509 -noout -enddate 2>/dev/null | cut -d= -f2) || return 0
[[ -n $end_date ]] || return 0
end_epoch=$(date -d "$end_date" +%s 2>/dev/null || date -j -f '%b %d %T %Y %Z' "$end_date" +%s)
remaining=$(( (end_epoch - $(date +%s)) / 86400 ))
(( remaining > days )) && return 0
echo "TLS certificate for ${host} expires in ${remaining} days"
return 1
}
Every message names the resource, the measurement, the threshold, and the top contributors. That is the difference between an alert someone acts on and one they acknowledge and forget.
Step 5: The Runner
run_check() {
local name=$1; shift
local output rc=0
output=$("$@" 2>&1) || rc=$?
if (( rc == 0 )); then
if record_success "$name"; then
notify resolved "${name} recovered" "${name} is healthy again."
info "RECOVERED ${name}"
else
info "OK ${name}"
fi
return 0
fi
if record_failure "$name"; then
if should_alert "$name"; then
notify critical "$name" "$output"
warn "ALERTED ${name}: ${output}"
else
info "SUPPRESSED ${name} (cooldown active)"
fi
else
local n; n=$(cat "${STATE_DIR}/fail-$(_key "$name")")
info "FAILING ${name} (${n}/${FAIL_THRESHOLD}) — damping"
fi
return 1
}
main() {
exec 200>"/var/lock/${SCRIPT_NAME%.sh}.lock"
flock -n 200 || { info "already running; skipping"; exit 0; }
local failures=0
run_check "disk" check_disk || (( failures++ )) || true
run_check "memory" check_memory || (( failures++ )) || true
run_check "svc-nginx" check_service nginx || (( failures++ )) || true
run_check "svc-postgres" check_service postgresql || (( failures++ )) || true
run_check "http-app" check_http "http://localhost:8080/health" || (( failures++ )) || true
run_check "cert-www" check_cert_expiry "www.example.com" || (( failures++ )) || true
# Dead-man switch: ping only on a completed run
if [[ -n ${HEALTHCHECK_URL:-} ]]; then
curl -fsS --max-time 10 "$HEALTHCHECK_URL" >/dev/null || warn "heartbeat failed"
fi
info "run complete: ${failures} checks failing"
return 0 # ALWAYS 0 — cron mail is not the alerting channel
}
if [[ "${BASH_SOURCE[0]}" == "${0}" ]]; then
main "$@"
fi
Two design decisions worth defending:
return 0 regardless of failures. Alerts go through notify. If the script also exits non-zero, cron mails you a second copy of everything, and people start filtering cron mail — including the message telling them the monitor itself broke.
The heartbeat only fires at the end. A dead-man switch that pings at the start cannot detect a run that hangs halfway.
Step 6: Deploy It
# /etc/systemd/system/monitor.service
[Unit]
Description=Health monitor
After=network-online.target
[Service]
Type=oneshot
EnvironmentFile=/etc/monitor/env # secrets live here, mode 0600
ExecStart=/usr/local/bin/monitor.sh
TimeoutStartSec=120
# /etc/systemd/system/monitor.timer
[Unit]
Description=Run health monitor every 5 minutes
[Timer]
OnBootSec=2min
OnUnitActiveSec=5min
RandomizedDelaySec=30
[Install]
WantedBy=timers.target
sudo install -m 0600 /dev/stdin /etc/monitor/env <<'EOF'
SLACK_WEBHOOK_URL=https://hooks.slack.com/services/...
HEALTHCHECK_URL=https://hc-ping.com/...
DISK_THRESHOLD=85
EOF
sudo systemctl enable --now monitor.timer
journalctl -u monitor.service -f
OnUnitActiveSec rather than OnCalendar means the next run is scheduled 5 minutes after the last one finished, which cannot overlap. EnvironmentFile with mode 0600 keeps the webhook out of the unit file and out of ps.
Verifying Your Work
# 1. Damping — one failure must be silent
MONITOR_DRY_RUN=1 DISK_THRESHOLD=0 ./monitor.sh # "FAILING disk (1/3)"
MONITOR_DRY_RUN=1 DISK_THRESHOLD=0 ./monitor.sh # "FAILING disk (2/3)"
MONITOR_DRY_RUN=1 DISK_THRESHOLD=0 ./monitor.sh # "ALERTED disk"
# 2. Deduplication — the 4th run must be suppressed
MONITOR_DRY_RUN=1 DISK_THRESHOLD=0 ./monitor.sh # "SUPPRESSED disk"
# 3. Recovery
MONITOR_DRY_RUN=1 ./monitor.sh # "RECOVERED disk"
# 4. Locking
./monitor.sh & ./monitor.sh; wait # second: "already running"
# 5. State inspection
ls -la /var/lib/monitor/
Extensions worth building: emit Prometheus textfile-collector metrics alongside the alerts so you get graphs for free; add a --check NAME flag to run one check; add severity tiers so disk-at-95% pages while disk-at-85% only posts to Slack; write the check functions into separate files under checks.d/ and source them in a loop, so adding a check is adding a file.
The lesson to take away: the checks are the easy half. Everything that makes this usable — damping, dedup, recovery, exit-zero — exists to protect the humans receiving the alerts. An alerting channel people mute is worse than no monitoring at all, because it creates the belief that something is watching.
Common mistake: Alerting on every failed check. One flapping service produces four hundred notifications overnight, everybody mutes the channel, and the monitor is now worse than nothing. Damping and deduplication are not polish — they are what makes the alerts readable, and they belong in the first version rather than the second.