Googlebot renders JavaScript before indexing a page. Most AI crawlers do not - they fetch the HTML and read what comes back. A single-page app therefore serves a complete page to a visitor and an empty shell to the crawler feeding ChatGPT, Claude or Perplexity. You will not see this in analytics, because the crawler never runs your JavaScript and never fires a pageview. The only reliable check is to fetch your own page the way they do: curl it, strip the tags, and read what is left.
Why can't AI crawlers see some websites?
Because most AI crawlers do not run JavaScript. GPTBot, ClaudeBot, PerplexityBot and CCBot read the HTML your server returns and move on, so a client-rendered page that builds its content in the browser looks empty to them. Googlebot renders JavaScript, which is why the problem rarely shows up in search rankings.
Googlebot has rendered JavaScript for years. It fetches your HTML, queues the page, runs the bundle in a headless browser, and indexes whatever appeared. That is expensive, and Google does it because search results are the product.
The crawlers building AI answer indexes mostly do not. They request a URL, read the bytes that come back, and move on. For a server-rendered site that changes nothing. For a client-rendered one it changes everything: the visitor sees your page, and the crawler sees a div and a script tag.
You will not find this in analytics. The crawler never executes your JavaScript, so it never fires a pageview, so the problem produces no signal at all in the place you would normally look.
How do you check whether AI crawlers can read your page?
The check is the same thing the crawler does. Fetch the page without a browser, strip the markup, and count what is left:
curl -s https://yoursite.com/your-page \ | sed -e 's/<script[^>]*>.*<\/script>//g' -e 's/<[^>]*>/ /g' \ | tr -s ' \n' ' ' | head -c 500
If that prints your headline and opening paragraph, you are fine. If it prints nothing, or a cookie banner and a navigation menu, that is your page as far as an AI assistant is concerned.
The numbers worth knowing: under 500 characters of visible text and a text-to-HTML ratio under 5% together mean the content is not in the HTML. Either signal alone is inconclusive - a markup-heavy page can still be perfectly readable.
Worth running against your robots.txt at the same time: a crawler that is disallowed cannot cite you however readable the page is. Check GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot by name - a group naming the agent beats the wildcard group, so a blanket Disallow: / under * is not the whole story.
What are the most common reasons AI crawlers miss content?
Three failures account for most of what we find: body copy rendered in the browser instead of on the server, numbers that animate up from zero and so ship as zero in the HTML, and every page inheriting one site-wide title and description. All three are invisible to a person using the site and obvious to a crawler.
1. Body copy rendered in the browser
The big one. A React or Vue application that fetches content client side ships an empty shell. The fix is a rendering change: Server Components or static generation in Next.js, and the equivalent in whatever you use. You rarely need it everywhere - the marketing pages, documentation and anything you want cited are enough, and the authenticated application can stay client-rendered. We do this as a fixed-scope engagement; see AI integration for how we scope that kind of work.
2. Values that animate from zero
A counter that animates from 0 to 47% ships >0< in the initial DOM. Every machine reading your page sees you claiming zero percent. It is a small bug with an unusually bad failure mode, and it is everywhere.
Put the true number in the markup and animate from it, not towards it: set state to the real value, drop to zero in an effect that only runs in the browser, then count back up. The server-rendered output and the no-JavaScript render both show the truth, and the animation still plays for everyone else.
3. Every page inheriting one title
A layout defines a title and description, no route overrides them, and every URL on the site ships identical metadata. It is one missing function, and it produces duplicate titles, duplicate descriptions and keyword-mismatch flags across the whole site at once.
It is also worth knowing how this looks in an audit tool: URL keyword-mismatch warnings often fire because the tool compares the slug against the title tag. Fix the titles and the URL warnings clear. Rewriting well-formed slugs to satisfy them throws away link equity to solve a problem that was never there.
How do you make a site readable to AI crawlers?
In this order: server-render the pages you want cited, allow the AI crawlers by name in robots.txt, give every route its own title, description and canonical, then put a complete answer in the first 100 words and use real tables for comparisons. An llms.txt file comes last, and is optional.
| Step | Effort | Why it matters |
|---|---|---|
| Server-render the pages you want cited | Days to weeks | Nothing else matters if the text is not in the HTML |
| Allow the AI crawlers by name in robots.txt | Minutes | A blocked crawler cannot cite you regardless of content quality |
| Give every route its own title, description and canonical | Hours | Duplicate metadata across a site is a sitewide finding from a single missing function |
| Put a complete answer in the first 100 words | Per page | That paragraph is what gets lifted into an answer |
| Use real tables for anything comparative | Per page | Tables get quoted; styled divs are invisible as structure |
| Add an llms.txt | Twenty minutes | Optional and low-value today. Do it last, and expect nothing. |
How do you stop the problem coming back?
Rendering regressions are silent. Someone converts a page to a client component for a legitimate reason, nothing breaks visibly, and the page quietly leaves the AI index. The only defence is an automated check in the deploy path.
Ours runs against the built output before anything ships and fails the deploy on: a route with no h1 or more than one, a title or description duplicated across routes, a missing canonical, JSON-LD that does not parse, a zero-state counter in the markup, and any route whose server-rendered HTML contains less than 250 words of text. That last check is the one that catches a rendering regression the day it happens rather than the quarter after.
- Run it on the build artifact, not on a dev server - the artifact is what you ship.
- Fail the build, do not warn. A warning in CI output is a warning nobody reads.
- Keep the checks in the repository next to the code, so relaxing one is a reviewed decision.
Does being readable get you cited by AI assistants?
Being readable is necessary, not sufficient. No amount of server-rendering makes an assistant choose your page over a better one, and nobody can currently measure AI citation share reliably - treat any tool claiming a precise number with suspicion.
What it does get you is a floor above zero. A page a crawler cannot read has a hard ceiling of never being cited, which makes this the first thing to check rather than the last. If you would rather hand the whole job over, our rates are published and this kind of remediation lands at the low end of them.
