AI MY SITE BLOG

Same site, wildly different AI search scores

The same small business can score in the high 70s on one AI assistant and the low teens on another. Here is why the gap happens and what closes it.

AI My Site·
Same site, wildly different AI search scores

The same small business can score in the high 70s on one AI assistant and the low teens on another, scanned on the same day. Across the 500 businesses in our Q2 2026 benchmark, more than one in five had a gap of 15 points or more between their strongest and weakest AI assistant, and one in twelve had a gap of 20 points or more.

So there is no single AI search score. This post covers why the same site reads so differently to ChatGPT, Claude and Gemini, what causes the gap, and what actually closes the weakest one.

One site, three different verdicts

When we scored 500 small businesses on ChatGPT, Claude and Gemini, the typical business had a gap of about 10 points between its best and worst assistant. That is the average. The tail is where it gets uncomfortable: more than one in five sites had a gap of 15 points or more, and 8.4% had a gap of 20 points or more.

Horizontal bar chart: the gap between a site's best and worst AI assistant, n=500. A gap of 20 or more points, 8.4% of sites. A gap of 15 or more points, 22.2%. A gap of 10 or more points, 50.2%. Share of the 500 small businesses we scanned.

The widest gap in our sample was about 66 points: a UK ecommerce site that one assistant rated in the high 70s while another placed it in the low teens. Visible and well recommended on one, almost absent on the other. It is not a rounding error, and it is not rare enough to ignore.

Why the same site reads so differently

AI assistants are not one audience. ChatGPT, Claude and Gemini look for different signals before they will name a business, so a site tuned for one can fall flat on another. In our sample Gemini tended to be the most willing to mention smaller brands, while ChatGPT and Claude asked for more evidence before recommending one. A business with strong category positioning might land well on the more generous assistant and get skipped by the stricter ones.

That is why a single site produces three verdicts. The content, the structure and the external mentions each carry different weight depending on which assistant is reading. Optimising for AI search in general, without knowing which assistant is weakest for you, quietly leaves the biggest gap untouched.

What closes the weakest assistant

The fixes are not different per assistant. The diagnosis is. The work that lifts a low score is the familiar list: clearer language about what you do and who you serve, structured data that maps to your business type, direct answers to the questions buyers actually ask, and the external mentions that give an assistant something to ground a recommendation in. What changes per assistant is which of these you are missing.

So the move is to find your lowest assistant first, then work the items that lift it, rather than spreading effort evenly across all three. That is the layer AI My Site is built to surface: your score on each assistant, then a single prioritised plan that lifts the weakest one. The action plan, not just the audit.

Common mistakes to avoid

  • Trusting one assistant's answer. Checking only the assistant you happen to use hides the gap. Your weakest one is where the lost recommendations are, and it may not be the one you looked at.
  • Optimising for AI search in general. A generic pass spreads effort evenly and leaves the biggest gap where it was. The lift comes from targeting the lowest score.
  • Assuming a strong score travels. Doing well on one assistant does not mean the others agree. In our sample the same site often scored 15 points apart or more across the three.
  • Chasing the assistant that already likes you. More effort on your strongest assistant adds little. The compounding gain sits in the one that barely mentions you.

How long does this take

Once you know which assistant is weakest and which items are dragging it down, the fixes are the ordinary ones, and anyone promising an overnight jump is guessing. In practice, businesses that work a prioritised list tend to see movement over a 4 to 12 week window, depending on how wide the gap is and how much they apply. What the data supports is narrower: the gap between assistants is real, it is often 15 points or more, and it is closeable with focused work rather than a rebuild.

Frequently asked questions

How can one site score so differently across AI assistants?

Because each assistant weighs different signals before it will name a business. The same content and structure can clear one assistant's bar and fall short of another's, which is why more than one in five sites in our sample scored 15 points or more apart.

Which assistant is usually the weak link?

It varies by site, so the honest answer is to check yours. Across the 500 we scanned, ChatGPT was the lowest-scoring assistant for 45.4% of businesses and Claude for 41.4%, so the weakest one is most often, but not always, one of those two.

Do I need a different plan for each assistant?

No. The fixes are largely the same across assistants. What differs is which ones you are missing, so a single prioritised plan aimed at your lowest assistant lifts visibility across all three from one set of changes.

How big is a typical gap?

About 10 points between a site's best and worst assistant, on average. Roughly half of sites sit at 10 points or more, more than one in five at 15 or more, and 8.4% at 20 or more.

See where you sit

If the same site can score 15 points apart across assistants, the useful question is which assistant is weakest for you and which items would lift it. We have published the full Q2 2026 State of AI Search benchmark, with the per-assistant spread and the breakdown by industry, country and platform, so you can see where your category sits.

Read the full Q2 2026 benchmark to see where you sit.

AI My Site Research · Q2 2026 · n=500. Spread is the gap between a site's highest and lowest of three AI assistants (ChatGPT, Claude, Gemini); Perplexity is excluded from headline figures. We measured AI search visibility, not revenue or conversion. Figures are a Q2 2026 point-in-time snapshot.

See how your own site scores

AI My Site checks all of this automatically and gives you a prioritised, plain-English plan for what to fix first.

Get your scores — £60/month