0 lines of javascript for a crawler: the 1,416 bytes of html bing reads instead of my react app
dev · Oct 4, 2026 · 6 min read
this site is a client-rendered react app. it ships one html file. there are 539 urls in the sitemap and every one of them is served that same document, because there is no /blog/index.html and no /it/index.html — there is exactly one html file in the build output and it is 12,245 bytes.
so what does bing see?
a crawler that does not execute javascript looks at that document and finds a perfectly good <head> and an empty <div id="root">. 10,377 bytes of metadata and zero words of content.
the fix was to stop shipping an empty div.
the fallback
33 lines of plain html inside #root, which react throws away on mount:
| measurement | value |
|---|---|
| lines | 33 |
| bytes | 1,416 |
| words | 101 |
<h1> / <h2> / <li> |
1 / 2 / 7 |
| links | 7 |
<script> tags |
0 |
one h1, a short paragraph that says who i am and what i do, a "services" list, a "contact" list, and four links into the site. 1,416 bytes of marketing copy that i maintain by hand, forever, for a reader that will never look at it.
plus 6 css rules — 13 lines, 543 bytes — so it arrives looking like a small page instead of a broken one. that part is not decoration. a fallback rendered as unstyled black-on-white text reads to a human reviewer as this site is broken, and the same 543 bytes are the difference between this site has minimal content and this site is broken.
why it is not a duplicate of the page
because createRoot clears its container. this is all the mount code does:
createRoot(document.getElementById('root')!).render(
<StrictMode>
<ThemeProvider>
<App />
react does not merge into #root, it wipes the children and puts its own. so the visitor never sees the fallback for more than a few milliseconds, and the crawler never sees react. there is no hydration mismatch to worry about and no "is the app ready yet" state to design around.
the compromise, stated plainly — and then fixed
i originally did not prerender. astro, or next with static export, or vite-ssg would have given every one of the urls its own html with its own body text. i did not take it, because it turns a marketing page into a second build pipeline with its own cache and its own deploy failure mode.
then i measured what the floor actually cost, and the answer was worse than i had assumed: 374 urls, and every one of them was the same document. the canonical on all 374 said the site root, which is a way of telling a crawler that the other 373 are duplicates of the home page. one document, 374 addresses.
so i wrote a prerenderer instead. it emits one html file per (route, locale) pair into the build output, generated from the same data the app reads, so the two cannot disagree. pnpm run build now writes 189 documents — 7 page routes × 7 locales, and one document per post per locale that post has been translated into. a language contributes a url the moment that one post is written, and the hreflang cluster for a post is computed from its own id, so the document and its alternates can never disagree about which languages exist.
here is the honest accounting, per url, now:
| what | route-specific? |
|---|---|
<title>, description, canonical, hreflang |
yes — statically in the head, one per url |
| JSON-LD | yes — BlogPosting and BreadcrumbList on posts, WebPage elsewhere |
| sitemap, rss | yes — 35 routes, split across the two locale clusters |
| body text | yes on posts and on the blog index — the whole article, in that language |
body text on /lab, /now, /lens |
partly. title, description and navigation, but not the prose, because that prose lives in jsx and there is no data file to read it from |
| the app | yes, one javascript bundle |
the interesting part was not writing the generator. it was discovering that the hreflang cluster had to be per kind of route. a post exists in three languages and a page in seven, so a single hardcoded list of eleven was advertising six languages that served the english original under a foreign-language tag — which google counts as duplicate content, not as localization.
both lists are now derived from the module files on disk. dropping posts/fr.ts in makes french appear, and nothing else has to change.
the floor is gone for the posts. it is still a floor for three pages, and i would rather say so than pretend otherwise.
the layer above the fallback, and the two bugs in it
the head is not the only thing that is route-specific. /blog/* has a cloudflare pages function that rewrites the metadata for known bots:
const BOT_AGENTS = [
"Twitterbot", "facebookexternalhit", "LinkedInBot", "Slackbot",
"TelegramBot", "WhatsApp", "Discordbot", "discordbot",
"ia_archiver", "Googlebot", "bingbot", "Applebot",
];
twelve user agents. for those, [id].js looks the post up in a generated og-data.js, then rewrites <title>, og:title, og:description, og:url, og:type, the twitter tags and the meta description through HTMLRewriter. bingbot is in the list, and bingbot is the one that does not run javascript. so for a blog post, bing gets the real title of the real post.
and then it gets the homepage's 101 words underneath it.
bug 1: DuckDuckDuckBot is not in that list. twelve entries, and the crawler I named in the comment directly above the fallback block — the whole reason the fallback exists — is the one that falls through to the defaults. every post on the site is titled Enea | Blog as far as duckduckgo is concerned. one string. that is the entire fix and i have not applied it.
bug 2: canonical is never rewritten. the map in MetaRewriter includes og:url but there is no handler for link[rel=canonical]. so the same document tells social platforms "this url is /blog/my-post" and tells search engines "this url is https://eneawork.it/". og:url and rel=canonical are not allowed to disagree, and mine do, on all 42 posts. the fix is one more entry in the map and one more .on("link[rel=canonical]") handler.
smaller: the function rewrites twitter:card to summary_large_image while og:image stays the 460×460 profile photo, because every post shares one image. a large card with a square avatar in it.
what i will not pretend about this
the fallback is marketing copy sitting in a code path. the moment the live site changes and i forget to change it, the crawler is being served a stale description of a different site, and nothing will fail. no test breaks, no type error, no build error. it is the kind of debt that only shows up in traffic you cannot attribute.
i have already left one small fossil in it: the line reads Fotografia: <a href="/lens">lens</a> — an italian word in an otherwise english document, in the one piece of html on this site that is supposed to be the permanent version of me. it is not a bug. it is what hand-maintained copy looks like after a year.
what i would tell my past self
- the head is where the seo lives, and csr does not touch it. every piece of metadata that matters is static, in the document, before any javascript runs. if you write the head by hand you have already done 80% of the seo work of a static site, and the remaining 20% is the thing people mean when they say "you should prerender".
- a fallback has to be the shortest thing that is still true. 1,416 bytes is a paragraph. 1,416 bytes of per-route content is 539 documents. the first one is a weekend; the second one is a platform.
- when you write a comment naming a threat, check that you actually handled it. i wrote "bing, duckduckgo and ai crawlers that do not run javascript" three lines above a fallback, and then shipped an allowlist of twelve bots with the second name missing. the comment was more thorough than the code. it usually is.