Quick one before we start: no AI Build Night tonight, I'm moving the boat. Back next Wednesday, 6pm London. Come along then.
My dashboard went red on a Thursday morning. One of my automations looked dead.
It wasn't. The job had run on schedule the previous two Sundays, done its work, and sent what it was supposed to send. The dashboard was wrong. Not the automation.
That was the first false alarm. By Tuesday I'd had three more, and I'd learned something I'd rather have learned from someone else's mistake.
How a healthy job turns a tile red
I run a control panel for my scheduled automations. Each one gets a tile. Green means fine, amber means look at it, red means it's broken. Simple.
The panel doesn't ask the job whether it worked. It reads a log file the job is supposed to leave behind. No fresh line in the log, no green tile.
So here's what happened across five days.
The weekly fitness check-in went red. It had run and done its work. It just never wrote its log line. The panel saw silence and assumed death.
My daily accounting sync went amber. It runs weekly in practice, and it ran fine on Monday. Same story: no log line, so the tile sulked. I tapped the check button twice and got the same answer twice, which at least told me the diagnosis was stable.
The YouTube Build Night uploader went amber too, and this one was funnier. It was an old cron from when I used Zoom recordings for uploads. Its gatekeeping code crashed every time it started. But the Build Night video had already been uploaded by a completely different pipeline that pulls from Fathom. The old job was a ghost, failing at a task someone else had already finished.
The newsletter monitor was the only real problem, and even that one wasn't the problem you'd guess. The newsletter itself went out fine on 1 and 5 October. The monitor job that checks it was getting killed by a ten minute script timeout before it could finish checking. The thing being watched was healthy. The watcher kept dying.
Four alarms. One real fault, and it was in the watcher.
What I changed
Each one got fixed at the source. I didn't just mute the alerts, because muting is how you end up with the next problem.
For the two missing log lines, I made the sync read the run history as a fallback, so a job that ran but forgot its paperwork still shows green. I also added an explicit log step to the accounting prompt, so it leaves a line from now on.
For the ghost uploader, I paused it. Not deleted. If I ever go back to Zoom uploads, it's sitting there. So right now it's grey on the panel, which is the honest colour.
For the monitor, I moved it to 08:30 UTC and capped it at eight minutes so it can't sit there strangling itself for forty.
But none of that was clever. And all of it was boring. That's the point.
What a good alert checks
So what should the panel look at instead? Here's where I've landed, and it applies to a plumber's quote follow-up as much as my setup.
First, did the thing happen to a real person? A text that reached a phone beats a log line that says "sent". Second, did something change on the other side? A booking that appeared in the calendar, a payment that landed, a reply that came back. And third, if neither is available, how long since this last did anything? A weekly job that's been silent for nine days is worth a look, whatever its paperwork says.
Look, none of this needs a fancy tool. A spreadsheet and a Monday morning habit would do. The tool isn't the hard bit. Deciding what "worked" actually means is the hard bit, and most people never write it down.
There's one more thing worth saying. The diagnosis was the slow part, because each alarm looked identical on the surface. Red is red. It doesn't tell you whether the job died or the paperwork did. So I now ask the same two questions every time: did the work happen, and did the record happen? Usually only one of them is true, and that tells you where to look.
The real cost of a cry-wolf dashboard
Here's what bothered me. Every false alarm costs more than the minutes it takes to diagnose.
After the second one, I caught myself glancing at a red tile and thinking "probably the log thing again." That's the dangerous moment. The day a job genuinely dies, I'll have trained myself to shrug at exactly the colour that should make me move.
An alert nobody believes is worse than no alert, because it feels like safety. You've got a monitor, so you must be covered. And meanwhile you've stopped listening to it.
If you run any kind of automation in your business, whether it's a missed-call text, an invoice chaser or a booking reminder, ask yourself one question. What does your alert actually check? The outcome, or the paperwork?
Checking the paperwork is easy. Did a log line appear? Did an email land in the sent folder? Did a status field flip to "done"? But a job can do all of that and achieve nothing. And a job can do its work perfectly and leave no paperwork at all. I just watched four of mine prove it in a week.
A better test for your own setup
So pick one automation you trust. Break it on purpose, in a safe way. Point it at a test contact, or switch off the thing it depends on for ten minutes.
Then wait, and see what tells you. If nothing does, you don't have monitoring. You have a feeling. If the wrong thing tells you, or it takes a day, you've got a leak.
And do the reverse too. Find the last alert you dismissed as "probably nothing." Work out why it fired. Either it was nothing, and the alert needs fixing, or it was something, and you got lucky.
I'd rather have a panel that's green when things are fine and red when they're not, and trust it when it says so. Mine's sort of closer to that this week than it was last week. That's the whole win.
Now, next Wednesday, I'll build one of these systems live. Alerts, dashboards, and the dull bits that keep an autonomous business honest. Bring your own false alarm and we'll see if we can break it on camera.
Next week's AI Build Night. 6pm London, live on Zoom, me building one of these systems from scratch while you watch. No slides, no rehearsed demo, and when it breaks on camera you get to see how I unbreak it. That's usually the useful bit. Grab a seat.
Brewed by Steven, poured by Viktor
About Steven Tann: Steven helps business owners build systems that run themselves using AI. After 10+ years helping 7,000+ businesses and building his own autonomous operations, he's the bloke who actually does it, not just talks about it. Find out more at steventann.com.