A Green Checkmark Is Not Evidence

My scrapers reported 100% success for weeks while inserting nothing. Here is how I changed monitoring to check what a pipeline produced, not whether it exited cleanly.

contents (7)
  1. What “success” was hiding
  2. Check outcomes, not exit codes
  3. Thresholds from real numbers
  4. Severity that means something
  5. The monitor can lie too
  6. Stale should look stale
  7. What I check now

Plus234Feed runs dozens of scheduled jobs: scrapers for Nigerian publishers, data pipelines for exchange rates and market data, embeddings, newsletters, push notifications. For a long time my monitoring answered one question: did the job succeed?

The dashboard said yes. It was wrong for weeks.

What “success” was hiding

Every one of these failures reported itself as a success at the time:

  • Two scrapers ran 24 times a day, returned HTTP 200 and inserted nothing for weeks. The dashboard showed a 100% success rate.
  • After a deploy, Google Analytics page views per user fell from 1.73 to 0.33. No error was raised anywhere. I noticed it by eye, days later.
  • Push notifications went out every night with a 0% click-through rate.
  • Some articles were published with advertising banners where their images should have been.

Nothing crashed. Every job exited cleanly. The checks were measuring whether code ran, not whether it worked.

Check outcomes, not exit codes

So I wrote a health job that looks at what came out the other end. It checks service reachability, scraper output, daily ingestion volume, the embedding backlog, how many recent articles have related stories, images, the newsletter, push volume and click-through, page views and search impressions.

The most useful check turned out to be the simplest. For scrapers, one distinction carries most of the weight:

0 inserted, 443 duplicates -> healthy: ran fine, nothing new to add
0 inserted, 0 duplicates -> broken: fetched nothing at all

Both show up as “0 inserted” on a dashboard, which is exactly why the broken scrapers went unnoticed. Duplicates are proof of life. They mean the scraper fetched real pages and the pipeline recognised articles it already had. A scraper that sees nothing at all, not even old articles, is not quiet. It is broken.

Thresholds from real numbers

A threshold that fires on a normal day trains everyone to ignore it. So every threshold was set against two weeks of production data, not picked out of the air:

Check Normal Alert when
Articles ingested per day about 728 below 40% of the 14-day average
Articles waiting for embeddings about 9 above 500
Articles published without an image 5.7% above 20%
Weekly newsletter every Friday nothing sent in 8 days

The embedding backlog is a good example of a silent failure. If the embedding job stops, nothing errors. The related-stories feature just gets quietly worse, falling back to “recent articles in the same category”, and no one notices.

Severity that means something

The job only fails, turning red, when something is critical, meaning the pipeline has stopped. A single dead scraper is degraded: it posts to Slack and the job passes. Slack only hears about a run when the status changes, plus one summary a day.

The goal is that red always means “act now”. If one broken scraper turned the job red every morning, red would stop meaning anything within a week.

The monitor can lie too

After I turned on bot protection for the website, the health check started reporting the site as down. It wasn’t. Readers were being served normally. The bot protection was challenging the health check itself: its requests came from GitHub Actions, which runs on datacenter IP ranges, and looked automated.

The fix was not just to get past the challenge. The check now recognises a challenge page and reports it as a challenge, rather than as “site down” or “site up”. A health check that silently reports the wrong thing in either direction is worse than no check.

Stale should look stale

The results feed a panel in SignalDesk, the internal portal I run the platform from. That panel only reads what the health job recorded. It never re-runs checks itself.

That raised a new question: what if the health job stops? The panel would keep showing the last result, which was green. So the endpoint marks a result as stale once it is more than 7 hours old. The job runs every 3 hours, so 7 hours means at least two runs were missed: tight enough to notice a stopped job, loose enough to survive one slow scheduled run.

A stale green panel is worse than no panel. It tells you everything is fine with full confidence and no evidence.

What I check now

When I add a scheduled job, I ask four questions:

  1. What does this job produce? Rows, emails, files, page views. Check that, not the exit code.
  2. What does “healthy but idle” look like, compared with “broken”? For scrapers it was duplicates. Every job has some equivalent.
  3. What is normal? Set thresholds from real data, so an alert means something changed.
  4. Who watches the watcher? If the monitor stops or gets blocked, it should say so instead of showing its last good answer.

None of this is complicated. The whole health check is a handful of database queries and a few HTTP requests, and it costs nothing to run. It just checks the right thing.