Insights

The five things that decide whether AI can find you.

This is the rubric we score every business against, explained well enough that you could check most of it yourself. Four of the five are things you control directly.

Why five, and why equal weight

Being recommended by an AI system is not one thing. It is a chain, and a business can break it at any link while doing everything else well.

A site can be beautifully written and unreachable by the crawlers behind AI answers. It can be reachable and structured so that no clean answer can be lifted out of it. It can be extractable and four years stale. It can be current and mentioned nowhere the systems trust. And all four can be fine while the answers still name somebody else.

So the score is composed of five pillars carrying equal weight, at 20% each. Equal weight is a deliberate choice rather than a default. Weighting toward the technical pillars would flatter agencies who only do technical work, and weighting toward visibility would flatter whoever happens to be winning already. A chain does not have a most important link.

1. Technical Readiness

The question: can AI crawlers reach your site, and is there anything there when they arrive?

The crawlers behind AI answers are not the same as Googlebot, they are not all as patient, and most of them largely do not run JavaScript. That last fact catches more businesses than anything else in this pillar.

What gets checked: whether AI crawlers are blocked in robots.txt, whether your CDN or firewall is blocking them at the edge without anyone knowing, whether your pages still contain their content with JavaScript switched off, server response codes and redirect chains, and whether the site is fast enough not to be abandoned mid-crawl.

What typically fails: a modern JavaScript site that renders its content in the browser. To a person it looks perfect. To a crawler that does not execute scripts it is a near-empty page. The business is not being outranked, it is being read as blank.

How to check it yourself: disable JavaScript in your browser and load your three most important pages. Whatever remains is roughly what a large share of these crawlers see. Then read your robots.txt and see which user agents you are refusing.

2. Structured Data and Extractability

The question: is there a machine-readable declaration of who you are, and is the page organised so an answer can be lifted out cleanly?

Two halves. The first is schema markup, the structured statement in your page’s code that says this is an organisation, this is its name, this is what it does, this is where it operates. The second is whether the visible content is shaped so a specific answer can be extracted from it.

What gets checked: whether Organization schema exists and is complete, whether it is consistent with the visible copy, whether service and article markup is present where it should be, heading structure, whether sections are self-contained, and whether questions on the page are answered directly rather than approached gradually.

What typically fails: the answer buried in the middle of a paragraph that builds up to it. AI systems extract from structure. A section that leads with its answer and then supports it can be quoted whole. A section that arrives at its answer in the fourth sentence usually cannot.

How to check it yourself: take any page and ask whether a stranger could lift one section out, read it alone, and get a complete answer. If they need the section above it for context, it will not be quoted.

3. Content and Freshness

The question: does your content answer real questions, in enough depth, recently enough to be cited?

This pillar is where the strongest mechanical fact in the discipline lives: 50% of AI citations come from content published or updated within the last 13 weeks (Amsive, 2025). Half the citations go to content less than a quarter old.

What gets checked: coverage of the questions your customers actually ask, depth against what is already being cited on those questions, publication and update dates, whether dates are exposed in the markup, and whether the site shows signs of an ongoing cadence or a burst of activity followed by silence.

What typically fails: a good site that stopped. Twenty solid pages, all published in the same eighteen-month period, none touched since. It reads to a person as a complete site and to these systems as a dormant one.

How to check it yourself: find the most recent substantive update to your site. If you have to think about it, that is the finding.

4. Authority and Entity

The question: do the third-party sources these systems trust actually mention you, consistently?

This is the pillar businesses have usually never worked on, and it carries more weight than most expect. Roughly 80% of AI citations come from third-party sources rather than a brand’s own content (BuzzStream, 12,000 AI responses, 2026). Wikipedia and Reddit alone drive over 25% of ChatGPT’s US citations (5W Research, 2025).

Your own site largely decides whether you are cited as the source and whether the description is accurate. Those other places largely decide whether you come up at all.

What gets checked: presence and consistency across directories, review platforms, industry roundups and community discussion; your Google Business Profile and its categories; whether your name, address, service list and description match character for character everywhere they appear; and whether the entity is unambiguous or gets confused with a similarly named business.

What typically fails: inconsistency rather than absence. Three versions of the address, two versions of the company name, a service list that differs between the site and the profile. Each is trivially fixable and together they teach these systems that the facts about you are uncertain, which is exactly when an answer starts hedging or inventing.

How to check it yourself: search your business name and open the first ten results that are not your own site. Read what each says about you. Note every detail that does not match your own site exactly.

5. AI Visibility

The question: when we ask the real questions, do you appear, and what is said about you?

The other four are inputs. This one is the outcome, and it is measured by asking rather than by inferring.

What gets checked: a locked set of the questions your customers actually ask, each run three times per surface because these systems do not answer identically twice, across ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Gemini and Copilot. From that: how often you are named, how often you are linked, whether what is said is accurate, how you are positioned against named competitors on identical questions, and which of your service categories the systems associate with you.

What typically fails: the surface split. A business appears reliably in one place and nowhere in another, and had assumed the two moved together. Google AI Overviews and Google AI Mode cite the same URLs only 13.7% of the time (Ahrefs, 540,000 query pairs, 2025), so being in one really is close to no evidence about the other.

How to check it yourself: partially. You can ask the questions and see the answers. What you cannot easily do alone is run a locked set consistently across six surfaces, three times each, and keep it comparable for a rerun in ninety days.

What a low score in one pillar actually means

One thing worth knowing before you read any AI visibility score, ours or anybody’s.

Most of the Authority pillar cannot be scored by a crawler. Seven of its nine checks live off your website entirely, in places no automated scan reaches. A pillar reading low on an automated report very often means nobody has looked yet, not that you scored badly.

We say which is which on every report, because letting “we did not check” quietly become “it is not there” is the easiest way to make a score look rigorous while meaning nothing. When you are comparing scores from different providers, ask which checks were automated and which were verified by a person. The answer separates the reports quickly.

Where to start

The order is not the order of the pillars. It is cheapest and most consequential first.

  1. Technical Readiness, because a site that cannot be read scores zero on everything downstream no matter how good it is. It is also usually a configuration fix rather than a project.
  2. Authority and Entity consistency, because matching your details everywhere costs almost nothing and directly reduces the chance of a wrong answer being given about you.
  3. Structured Data, because it is a one-time build that keeps paying.
  4. Content and Freshness, because it is the slowest and the only one that requires an ongoing commitment rather than a fix.
  5. AI Visibility, which is not a thing you do. It is the thing that moves when the other four improve, which is why it is measured before and after rather than worked on directly.

All five pillars scored, every client-facing finding verified by a person against the evidence, and a fix list where each item carries its business consequence and its owner.