backend
article··7 min read

Cloudflare cached HTML instead of CSS: notes from moving a static site

#cloudflare #apache #caching #deployment #htaccess

This blog is a static Nuxt site on shared hosting, behind Cloudflare. Today I moved it to a new directory on the same account: the built site is uploaded by a script over FTPS, and the domain is switched to the new directory in the control panel. After the switch the home page rendered as white text with no styles, while every check from the terminal said everything was fine. Below is where that mismatch came from and what changed so it does not happen again.

Symptom: curl says 200, the browser does not load CSS

The first check looked good. The HTML referenced /_nuxt/entry.Czy_6vul.css and that file answered correctly:

curl -s -o /dev/null -w '%{http_code} %{content_type}\n' https://tkulesza.eu/_nuxt/entry.Czy_6vul.css
# 200 text/css

A headless browser (puppeteer) showed what was actually happening. The console had this error:

Failed to load module script: Expected a JavaScript-or-Wasm module script
but the server responded with a MIME type of "text/html".

After adding a listener that logs _nuxt responses whose Content-Type contains html, it turned out that the same CSS URL and several JS modules reached the browser as HTML, with status 200:

page.on('response', r => {
  const ct = r.headers()['content-type'] || ''
  if (r.url().includes('/_nuxt/') && ct.includes('html')) {
    console.log('BAD', r.status(), ct, r.url())
  }
})

The status was right and the type was wrong. A check that looks only at the status code will not catch this.

Cause: a 200 fallback and caching by extension

Two things combined.

The first is the server. On this host a request for a file that does not exist did not end in a 404. It returned the home page with status 200. The setting comes from configuration outside my directory, so I had not seen it before. For a single-page application this is sometimes intended. For a static site with prerendered files it is harmful.

The second is Cloudflare. By default it caches responses based on the extension in the URL (.css, .js, images), not on the content type the origin returned. If the origin answers entry.Czy_6vul.css with HTML and status 200, Cloudflare stores that HTML under that URL and keeps serving it.

During the move there was a window in which the new HTML with new file names was already being served, while requests for those files still reached the old directory. The old directory did not have them, so it answered with the home page, and Cloudflare stored it. Hashed file names make this worse: the content does not change, so the name does not change, and the bad response stays in cache for as long as the TTL allows.

Why did curl get the correct file? The browser and curl sent different headers (the browser asks for compression, among other things) and hit different cache entries. Both responses had cf-cache-status: HIT. I did not dig into which header separates the entries, because purging the cache fixed the problem. The practical conclusion is simpler: a curl check does not replace a browser check.

Fix

The immediate fix was to purge the cache in the Cloudflare dashboard (Caching, Configuration, Purge Everything). To keep it from coming back on the next deploy, the site now ships its own .htaccess, in which a missing file ends in a 404:

DirectorySlash Off
DirectoryIndex index.html
ErrorDocument 404 /404.html

RewriteEngine On

# /x/ → /x when a prerendered page exists
RewriteCond %{DOCUMENT_ROOT}/$1/index.html -f
RewriteRule ^(.+)/$ https://tkulesza.eu/$1 [R=301,L,NE]

# /x → /x/index.html without a redirect
RewriteCond %{DOCUMENT_ROOT}/$1/index.html -f
RewriteRule ^(.+)$ $1/index.html [L]

# Anything else that is not on disk → 404
RewriteCond %{REQUEST_FILENAME} !-f
RewriteCond %{REQUEST_FILENAME} !-d
RewriteRule ^ - [R=404,L]

I first tried FallbackResource disabled, assuming the fallback came from that directive in a parent directory. It changed nothing, so the mechanism is something else. An explicit 404 rule at the end works regardless of what the host has configured.

Hashed files also got a header that lets browsers and the CDN keep them for a long time. The exception is _nuxt/builds/, where Nuxt keeps a file with the current build id:

<If "%{REQUEST_URI} =~ m#^/_nuxt/# && %{REQUEST_URI} !~ m#^/_nuxt/builds/#">
  Header always set Cache-Control "public, max-age=31536000, immutable"
</If>
<Else>
  Header always set Cache-Control "no-cache"
</Else>

Upload order matters too. The deploy script now uploads the hashed files first and the HTML that references them second, so there is no moment in which the HTML points at a file that is not there yet.

Second issue: allowing traffic only through Cloudflare

While at it, I closed direct access to the origin. Anyone who knows the host's IP can bypass Cloudflare by sending a request with the domain's Host header straight to that address. The usual answer is to allow only the addresses on Cloudflare's list, for example with Require ip.

Before uploading such a rule I checked which address the server sees. A temporary PHP script printed REMOTE_ADDR for a request made through Cloudflare:

<?php
header('Content-Type: text/plain');
echo $_SERVER['REMOTE_ADDR'] ?? '-', ' | ', $_SERVER['HTTP_CF_CONNECTING_IP'] ?? '-';

It printed the same address twice: mine, not Cloudflare's. The host rewrites the client address (in Apache this is mod_remoteip), so REMOTE_ADDR already holds the visitor's IP. A Require ip rule with Cloudflare's ranges would have blocked every reader and let through only those whose own address happens to be in a Cloudflare range.

Apache exposes CONN_REMOTE_ADDR in expressions: the address of the actual TCP connection, which mod_remoteip does not touch. The rule checks that variable:

RewriteCond expr "!(%{CONN_REMOTE_ADDR} -ipmatch '173.245.48.0/20' || %{CONN_REMOTE_ADDR} -ipmatch '103.21.244.0/22' || ...)"
RewriteRule ^ - [F,L]

The list has 15 IPv4 and 7 IPv6 ranges. After uploading I checked both paths:

curl -s -o /dev/null -w '%{http_code}\n' https://tkulesza.eu/
# 200
curl -sk --resolve tkulesza.eu:443:<host IP> -o /dev/null -w '%{http_code}\n' https://tkulesza.eu/
# 403

The PHP probe was deleted right after the check. Cloudflare's address list changes rarely, but it does change. If the site ever starts returning 403 through Cloudflare, an outdated list is the first suspect.

A note on scanning your own site

While checking whether files such as .git/config or .env could be fetched from outside, I also requested a few common paths, wp-login.php among them. The host's bot protection treated that as an attack and for several minutes served my address a "One moment, please..." page instead of the site. I did the rest of the review on the built files locally. If your host has this kind of protection and your home network and server share an outgoing address, keep it in mind.

Checklist for moving a static site behind a CDN

  • A missing file must return 404, not the home page with status 200. Test it on a URL that certainly does not exist.
  • Check Content-Type, not only the status code. A stylesheet served as text/html still has status 200.
  • Test in a real browser or with puppeteer. Curl sends different headers and can hit a different cache entry.
  • After switching directory or server, purge the CDN cache before calling the deploy done.
  • Upload hashed files before the HTML that references them.
  • Before restricting access to CDN addresses, check whether REMOTE_ADDR has already been rewritten to the client address.
end of node