Edge Caching in Front of a Python API: Lessons From Production

A Cloudflare Worker cache in front of a FastAPI backend: TTLs per path, never caching per-user data, serving stale copies during outages, and why every bad TTL I set showed up as a user-facing bug.

contents (7)
  1. The shape of the Worker
  2. One TTL per path
  3. Every wrong TTL became a bug report
  4. Never cache per-user data
  5. Serve a stale copy when the backend is down
  6. Purges are by exact key
  7. The checklist

Plus234Feed has a FastAPI backend with a couple of hundred endpoints: articles, search, exchange rates, stock market data, football, legislators, elections. In front of it sits a Cloudflare Worker that caches responses at the edge.

The idea is simple. Most readers ask for the same things, so answer them from a cache close to them instead of calling the backend every time. The details are where it goes wrong, and almost every mistake I made showed up as a bug a reader could see, not as a slow request.

The shape of the Worker

On each request, the Worker:

  1. Works out how long this path can be cached for.
  2. If the answer is zero, passes the request straight to the backend.
  3. Otherwise checks the cache, returns a hit, or fetches from the backend and stores the result.

The cache key is the full path plus the query string:

function getCacheKey(request) {
const url = new URL(request.url);
// The Cache API needs a fully qualified URL
return `https://cache.example.com${url.pathname}${url.search}`;
}

One TTL per path

Different data changes at very different speeds, so each path prefix gets its own cache lifetime in seconds:

const CACHE_TTL = {
'/v2/search': 60, // changes constantly
'/v2/exchange': 300, // rates move during the day
'/v2/articles/': 3600, // an article rarely changes once published
'/v2/legislators': 900,
'/v2/constituencies': 86400, // changes at a redistricting, not on a Tuesday
'/v2/elections/2027/results': 60, // must sit ABOVE the next line
'/v2/elections/2027': 21600,
default: 120,
};
function getCacheTTL(pathname) {
for (const path of NO_CACHE) {
if (pathname.startsWith(path) || pathname.includes(path)) return 0;
}
for (const [path, ttl] of Object.entries(CACHE_TTL)) {
if (path !== 'default' && pathname.startsWith(path)) return ttl;
}
return CACHE_TTL.default;
}

The first matching prefix wins, so order matters. Election results change by the minute on election night, so /v2/elections/2027/results must come before the broader /v2/elections/2027 entry. Put them the other way round and live results get cached for six hours.

Every wrong TTL became a bug report

The TTL table looks like configuration. In practice, each line is there because the previous value caused a visible problem.

Too long: the fix that didn’t show. The legislators endpoint once had a 24-hour TTL, on the reasoning that the roster only changes at elections. But the roster is re-imported every morning and corrected by hand when a seat changes. A stale cached response outlived a data fix: it predated a migration, had no image field at all, and every legislator’s portrait on the site fell back to initials. It is now 15 minutes.

Too short: paying for nothing. The constituency register (districts, constituencies and which local governments make up each) was falling through to the 2-minute default. A page was re-fetching every state’s seats every two minutes to be told the same thing. That data changes when boundaries are redrawn, so it now caches for a day, and the import job purges it when it does change.

The default is a decision too. Election candidate lists also fell through to the 2-minute default, so nearly every visit paid a cold 1 to 1.4 second trip to the backend. They now cache for six hours, with a purge when a list changes.

The lesson: anything on the default TTL has not been thought about yet. It is worth going through every endpoint and asking how often its data really changes.

Never cache per-user data

This is the rule that matters most, because breaking it is a privacy problem, not a performance problem.

// Paths that must never be cached: per-user data, writes, tracking.
const NO_CACHE = [
'/v2/auth/', '/my-answer', '/my-vote', '/my-answers',
'/v2/questions/vote', '/v2/questions/answer',
'/v2/articles/view', '/v2/articles/click',
'/user-engagement',
'/v2/newsletter/status',
];

Two of those entries were added after bugs.

A reader’s saved articles were falling through to the default cache. Saving an article and then opening the Saved page showed the list from before the save, so the feature looked broken. The cache key included the user ID, so nobody saw anyone else’s list, but per-user data that changes on every action has no business in a shared cache.

The newsletter subscription status had the same problem. Turning a newsletter on and then reading the setting back within two minutes returned the old value, so the switch appeared to undo itself.

Both were caught as “this feature is broken”, not as “this is cached”. When a write seems not to stick, check the cache before the database.

Serve a stale copy when the backend is down

When the Worker stores a fresh response, it also stores a second, longer-lived copy under a different key:

const staleTTL = Math.max(cacheTTL * 6, 3600); // 6x the TTL, at least an hour
const staleKey = getCacheKey(request).replace('cache.', 'stale.');
ctx.waitUntil(Promise.all([
cache.put(cacheKey, freshResponse),
cache.put(new Request(staleKey), staleResponse),
]));

If the backend later returns a 5xx, the Worker serves the stale copy instead of the error, and marks it with an X-Cache: STALE header so it is visible when debugging:

if (response.status >= 500) {
const stale = await cache.match(new Request(staleKey));
if (stale) {
const headers = new Headers(stale.headers);
headers.set('X-Cache', 'STALE');
headers.set('X-Backend-Status', String(response.status));
response = new Response(stale.body, { status: 200, headers });
}
}

A news site showing a page that is an hour old is far better than showing an error. The backend runs on a single hosted service, so this is cheap insurance against its bad minutes.

One detail worth getting right: copy headers with new Headers(stale.headers). Spreading a Headers object into a plain object ({ ...stale.headers }) gives you an empty object, because its entries are not ordinary properties, and the response quietly loses its Content-Type.

Purges are by exact key

When data changes before its TTL is up, the import job tells the Worker to purge it. The Cache API deletes by exact key, and the key includes the query string. So these are two different cache entries:

/v2/legislators?state=Lagos&chamber=senate&limit=100
/v2/legislators?state=Lagos&chamber=senate&limit=120

That is not academic. A roster correction was once invisible on the site for a day because only one of those entries was being purged, and the page was reading the other.

There is no prefix purge. A caller that wants a whole family of entries gone has to list them. So the job that writes the roster sends the exact URLs for the state it just changed, rather than something like /v2/legislators/*.

The practical rule: whoever writes the data should know exactly which URLs read it, and purge those.

The checklist

  1. Give every path a TTL based on how often its data really changes. Treat the default as “not decided yet”.
  2. Order prefix rules from most specific to least specific.
  3. Keep a never-cache list for anything per-user, any write, and any tracking endpoint.
  4. Store a long-lived stale copy and serve it when the backend fails.
  5. Make purges exact, and have the writer purge the URLs it affects.
  6. Mark responses with a cache status header, so HIT, MISS and STALE are visible when something looks wrong.

A cache in front of an API is a few hundred lines of code. Most of the engineering is in the table of numbers, and in knowing which endpoints should never be in it.