crawler-snapshot
Find out what AI and search crawlers really see on your JavaScript site, then save real HTML for them using the Chrome you already have.
- Status
- Live
- Launched
- Licence
- MIT
- Language
- JavaScript
One command, copied from the repo.
Install crawler-snapshot: npx crawler-snapshot check https://your-site.example --crawl
reidifydesign/crawler-snapshot
What it does
The check command fetches each route the way a non-rendering crawler does, with one request, a crawler User-Agent and no JavaScript. It then renders the same URL in your browser and compares the two: visible text, headings, links, JSON-LD blocks, the title, meta description, canonical link and robots meta tag, and content that sits in the markup but stays hidden until a script runs. Each route gets OK, WARN or FAIL, and each finding says what it saw. The exit code is 1 when anything fails, so it works as a CI step.
The snapshot command renders routes in your browser and writes out/<route>/index.html plus a manifest.json. It is resumable. Serve examples for Node, Express, a Netlify Edge Function and nginx send the snapshot to a listed crawler user agent and your normal app to everyone else. It is free, self-hosted and needs no account. It drives the Chrome, Edge or Chromium you already have, downloads nothing, and needs Node 18.3 or newer.
It is a command line tool you run yourself, not a hosted service. The nginx example has not been run against nginx, so test it before you rely on it. Server rendering is better where you can have it.
Why it exists
On 20 August 2026 we found that all 123 routes of reidify.design were serving their whole page inside a hidden div after the footer. It had been that way for months. The cause was one Suspense boundary above the route tree. When the site was prerendered, React took its streaming path and wrote the real page into a hidden div, and the small inline script that moves it into place never made it into the built files, because the site's Content Security Policy only allows scripts from itself plus a few hashes. Browsers never showed the problem, because React rebuilds the page when it mounts. Only readers of the served HTML were affected.
A check that counted characters did not catch it, because characters inside a hidden div are still characters. The check that did catch it read document structure. crawler-snapshot check is that second kind of check, made general.