What tells you a backup stopped running? A failure alert never fires for that

The restore test thread on here got me thinking about the step before it. Most backup alerting I've seen fires when the job fails. A cron line that quietly stopped firing, or a NAS mount that vanished, never fails anything. The job just doesn't run and the dashboard stays green. My setup is a nightly pg_dump to object storage on a systemd timer. The script writes its last success time as a Prometheus textfile metric, and the alert fires when that's older than 26 hours. Once a week a restore drill loads the newest dump into a throwaway Postgres and compares tables and row counts. The part that surprised me was a hole in the stale check itself. If node-exporter gets recreated without the textfile mount, the metric is just gone, and a rule comparing a timestamp that doesn't exist never fires. So now it also alerts when the metric is missing entirely. What do you use to catch the job that just stops? Healthchecks style pings, Uptime Kuma push monitors, something else?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论