Scaling a Security Scanner on Cloudflare Workers
How WebSentry uses a producer and consumer queue, R2 and an atomic claim to run scheduled scans reliably as the number of monitors grows.
contents (6)
WebSentry grades a website’s security in seconds: TLS, headers, content security policy, cookies, DNS and email authentication, vulnerable JavaScript libraries and more. It also monitors sites on a schedule and sends an alert when a grade drops.
It runs entirely on Cloudflare: Workers for compute, D1 for the database, R2 for storage, Queues for background jobs and Cron Triggers for scheduling. This post covers how the scheduled monitoring is built to scale and the bugs that shaped it.
The stack
| Layer | Service | Job |
|---|---|---|
| Compute | Workers | Every HTTP request and every scan |
| Database | D1 | Users, scans, monitors, rate limits |
| Blob storage | R2 | Full scan results as JSON |
| Job queue | Queues | Scheduled monitor scans |
| Scheduler | Cron Triggers | Starts the daily monitor run |
Producer and consumer
The naive version of scheduled monitoring is one cron job that loops over every monitor and scans each site. It works for ten monitors. At a few hundred, one invocation is doing hours of network requests, and a single slow or failing site holds up everything behind it.
So the work is split in two.
The cron job is the producer. It does almost nothing: it selects the active monitors that are due today (daily ones every day, weekly ones on Mondays, monthly ones on the 1st) and puts one message per monitor on a queue. Cloudflare Queues caps a batch send at 100 messages, so it sends in chunks of 100. Missing that limit is a silent failure, which is the worst kind.
The queue consumer does the work. Each consumer invocation receives a batch of up to 25 jobs, runs each scan, saves the results and sends any alerts. A job that throws is retried, up to twice, without affecting the others in the batch.
The batch size started at 5. With more than 100 monitors, the queue took many cycles to drain. Raising it to 25 made draining about five times faster.
This split means the cron job finishes in moments no matter how many monitors there are. Scans happen at whatever rate the queue delivers them, failures are isolated to one monitor, and retries come for free.
At-least-once delivery means duplicates
Both Cron Triggers and Queues guarantee at-least-once delivery. That is the right guarantee for reliability, and it has a consequence: sometimes a message arrives twice.
The first version of the job had no protection against that. Users started seeing two scans per monitor per run.
The first fix checked when the monitor last ran and skipped the job if it was too recent: 12 hours for a daily monitor, 60 for weekly, 240 for monthly. That helped, but left a race. Two deliveries of the same message could read the same old timestamp at the same moment and both decide to scan.
The real fix is to claim the monitor atomically before scanning:
UPDATE monitorsSET lock_until = ? -- now + 30 minutesWHERE id = ? AND is_active = 1 AND (lock_until IS NULL OR lock_until < datetime('now')) AND (last_scan_at IS NULL OR last_scan_at <= ?)Only one delivery can make that update succeed. The other sees zero changed rows and exits. The 30-minute lock also expires on its own, so a scan that crashes halfway does not block the monitor forever.
If your system uses queues, assume every message can arrive twice and make the handler safe to run twice. Idempotency is cheaper to design in than to debug later.
Keep the database lean with R2
A full scan result is a large JSON document: every check, every finding, every header. The first version stored it in a column in D1.
That was a problem waiting to happen. D1 is SQLite with a single write primary: reads are replicated and fast, but every write goes through one node. Large rows make every write and every read of the scans table heavier.
So the full results moved to R2, stored as one object per scan. D1 keeps only the small row: the URL, the grade, the time, the user. That keeps the table fast to query and the writes small.
The R2 write is also designed to fail safely. The scan metadata is saved to D1 first. If the R2 write fails, the scan still succeeds for the user and only the detailed report is unavailable. Reading works the same way: try R2 first, and fall back to the old D1 column for scans saved before the move. There was no big migration day; old and new rows simply coexist.
Small fixes that add up
Not every scaling problem is architectural. Two others came from data that only grew:
- Rate-limit rows were never deleted. Old rows outlived the window they were counted in, so the rate-limit query slowed down over time. The daily cron now deletes expired rows, which keeps the table a stable size.
- Indexes for the queries that actually run. Scans by user and date, monitors by status and frequency, monitor logs by monitor and time, and so on. Cheap to add, and the difference grows with the data.
Plan the next fix before you need it
The project keeps a scaling notes file. Each known limit has a signal to watch for and a fix already worked out, and the fixes above all started there. Writing these down before they are needed has been one of the most useful habits on this project: when a limit arrives, it is a planned change, not an emergency.