Answers

What AI agents look for on a website: structure they can parse without guessing

AI agents look for clean semantic HTML, a clear heading hierarchy, structured data and plain factual statements about the business, all of it readable without executing JavaScript. They are not evaluating how a page looks. They are extracting what a page says and how confidently it says it.

By Rish Sadh, founderUpdated

Short answer

An AI agent reads the underlying document, not the rendered page a person sees. It looks for semantic tags that say what each part of the page is, a heading tree it can follow in order, structured data that names the business and what it does, and factual sentences it can lift accurately. Visual polish is invisible to it either way.

Why a machine 'looks' at a page differently to a person

A person scans a page visually: the hero grabs attention, the layout implies hierarchy, colour and spacing signal what matters. An AI agent has none of that available to it in the same way. It works from the document's actual code: the tags, the order they appear in, and whatever text and data live inside them.

That gap matters most when a page depends on JavaScript to render its content. Many AI crawlers fetch a page's scripts but never run them, so whatever text or structure only exists after the page hydrates simply is not there from the agent's point of view, no matter how it looks to a visitor.

What agents are actually extracting

Three things, mostly. First, entity facts: what the business is called, what category it belongs to, where it operates and who runs it, usually pulled from structured data (schema.org markup) rather than inferred from prose. Second, the heading hierarchy: an h1 followed by a logical run of h2s tells an agent what the page is about and how its claims are organised, the same way a table of contents helps a person skim.

Third, plain factual statements. A sentence that states a claim directly, this business does X, is based in Y, works with clients in Z, is far easier for an agent to extract and repeat accurately than the same claim buried in a metaphor or a marketing headline built to be read aloud rather than parsed.

What agents effectively skip

Anything purely decorative: animation, background video, imagery with no accompanying text, and layout choices that exist to guide the eye rather than to state a fact. None of that is wasted on the human reader, but none of it registers with the agent either, so it should never be the only place a claim lives.

The same applies to content gated behind an interaction, a tab that only reveals its text on click, or copy that only appears after a hover or a scroll trigger. If a fact only exists inside an interaction, an agent that cannot perform that interaction never sees it.

A six-point check against what agents read

  1. The words are in the HTML before JavaScript runs

    View the page source, or load it with JavaScript switched off. If the headline, the services and the facts are not there, many AI crawlers will not see them either.

  2. One h1, then a logical run of h2s

    Read the headings alone, in order. They should work as a table of contents for the page's argument.

  3. Structured data names the business

    Organization or LocalBusiness markup on the homepage stating the name, category, location and who runs it, so none of it has to be inferred.

  4. The key facts are plain sentences

    What the business does, where, and for whom, each stated directly at least once, not only implied by a metaphor or a headline written to be read aloud.

  5. No claim lives only in an image or a video

    Animation, background footage and imagery can carry a claim for a person, but the same claim needs to exist as text somewhere on the page.

  6. Nothing important sits behind a click

    Text that appears only in a tab, on hover or after a scroll trigger may never be seen by an agent that cannot perform the interaction.

Two readers, one document

A page can be visually striking for a person and completely blank for the system reading its code underneath.

How this changes what 'good content' means

Writing for an agent does not mean writing worse copy for a person. It means making sure the same true, specific claims that already belong on the page also exist as plain text an agent can find: in a heading, in a short factual sentence, in structured data, rather than only implied by a photo or a layout choice.

In practice that is a semantic HTML decision as much as a copywriting one: the right tag for the right content, a heading tree that actually describes the page's argument, and structured data that names what the prose already says, so the two readers of the page, human and machine, are told the same thing in the way each of them can actually use.

They don't execute them. They can't read client-side rendered content.
Vercel, The rise of the AI crawler

23.84%

of Claude's crawler requests fetch JavaScript files, against 11.50% of ChatGPT's, measured across Vercel's network. Neither crawler executes them, so client-rendered content is not seen.

Vercel, AI crawler study

0.664

Correlation between branded web mentions and being named by AI, against 0.218 for backlinks across 75,000 brands. The authors caveat that correlation is not causation, and the sample skews to established domains.

Ahrefs, 75,000 brands

Negative

Measured effect of keyword stuffing on generative engine visibility. The tactic does not merely fail on these engines, it costs you.

GEO, KDD 2024

Common questions

Do AI agents care about how a website looks?

Not directly. An agent parses the underlying document, not the rendered visual layout, so colour, imagery and animation carry no weight for it. Those choices still matter for the human reader deciding whether to trust the business, which is why structure and design are handled as separate layers rather than traded off against each other.

Does JavaScript stop AI agents from reading a site?

It can. Vercel instrumented AI crawler traffic across its network and found that the crawlers behind ChatGPT and Claude request JavaScript files but do not execute them, on 11.50% and 23.84% of requests respectively. Anything that only exists after a page hydrates never reaches those agents, regardless of how it looks to a person.

Which schema markup matters most?

Whichever type actually describes what is on the page. Organization and LocalBusiness name the entity, Service or Product names what is offered, FAQPage marks up a visible question and answer set, and BreadcrumbList states where the page sits in the site. Adding a type for content the page does not actually contain does not help and can read as noise.

Can a page be optimized for agents without hurting the human experience?

Yes, and it usually should be. Semantic HTML, a clear heading hierarchy and structured data all sit underneath the visual design rather than replacing it. A page can be editorial and premium to look at while still being a clean, parseable document underneath.

Do AI agents read alt text on images?

Alt text is one of the few places an agent gets information about an image at all, since it cannot see the image itself the way a person does. A decorative image with no alt text is simply skipped; a product or proof photo with no alt text is a fact the page never actually states.

Next

Find out what an AI agent can actually read on your site today.