Most store owners have no idea whether ChatGPT can even see their business. This is the 10-minute check I run first on every store, including my mom's. No tools to buy. You need your website, an AI chat window, and a browser tab.
Crawler names and reports below were checked in July 2026. This part of the field changes fast, so verify against each provider's own documentation before you act on it.
Open ChatGPT and Perplexity. Ask five questions a real customer would ask, in plain words: your product category plus your town, a how-to question you know the answer to, and "what should I look for when buying" your main product. Watch for two things: does your business ever get named, and whose pages do the answers cite? Those cited pages are your real competition in AI search, and they are rarely who you think.
Do this in a fresh incognito window with no logins. A signed-in session carries your history and your past questions, which will show you a flattering answer no customer ever sees. Incognito does not clear your location, so a local business is still seeing a locally-weighted answer, but it removes the biggest source of self-flattery.
Type yourdomain.com/robots.txt into a browser and look for Disallow lines under the AI agents. I had this wrong myself until recently: the big providers run several agents for different jobs, and blocking one does not do what blocking another does.
For OpenAI, GPTBot is the training crawler, OAI-SearchBot is the one tied to appearing in ChatGPT's search results, and ChatGPT-User is the live fetch made while answering someone's question. OpenAI's documentation states these are independent, and gives allowing OAI-SearchBot while disallowing GPTBot as a worked example. So if your goal is to appear in ChatGPT and you have only checked GPTBot, you have checked the wrong line.
Anthropic splits the same way with ClaudeBot, Claude-User and Claude-SearchBot. Perplexity has PerplexityBot, and also Perplexity-User, which its own documentation says generally ignores robots.txt because a person asked for that specific fetch. Google-Extended governs Gemini apps and Vertex AI grounding, and Google states it does not affect inclusion in Google Search.
Run your best-selling product's page through Google's Rich Results Test (search the phrase, it's a free Google tool). It shows whether your structured data, the machine-readable description of your product, price, and stock, actually parses. On Shopify, your theme generates this automatically, but themes get it wrong, and the errors are easy to miss because nothing looks broken on the page. Errors here mean engines are guessing at your catalog.
Open your About page or a product page and read only the headings and first sentences. Could a stranger answer "what is this business, what does it sell, and where" from that skim? That skim is roughly what an AI extracts. If the answer lives in an image, a logo, or your head, the machine does not have it. My mom's shop had 990 product images with no alt text; the machine was blind to a third of the store.
Never named, but competitors are cited: you have a visibility gap and real upside; start with the entity and passage work in the plain-English GEO guide for small business owners. Blocked crawlers: unblock the right ones this week; it is the highest-leverage five-minute fix in this whole field. Schema errors: fixable in a weekend, and on Shopify usually a theme setting or an app away. Everything passing: you are ahead of most of your market, and the next fight is earning citations with content worth pulling.
Ask it directly, then check that you allowed in the right agent. Open ChatGPT in a fresh incognito window and ask three or four questions a customer would ask about your category and your town, and watch whether your business gets named at all. Then open yourdomain.com/robots.txt and look for Disallow lines under OAI-SearchBot and ChatGPT-User, not only GPTBot. Those two checks separate the two failure modes: the engine cannot reach you, or it can reach you and does not think you are worth citing. The fixes are completely different.
No, and this is the single most common mistake I see in robots.txt files. GPTBot is OpenAI's training crawler. Appearing in ChatGPT's search results is governed by OAI-SearchBot, and the live fetch made while answering a question comes from ChatGPT-User. OpenAI's own documentation says the settings are independent and uses exactly this example: you can allow OAI-SearchBot while disallowing GPTBot. Blocking GPTBot is a real decision about training data. It is not the switch for visibility.
More than one each, because the providers split them by job. OpenAI: GPTBot for training, OAI-SearchBot for search appearance, ChatGPT-User for answer-time fetches. Anthropic: ClaudeBot, Claude-User, Claude-SearchBot. Perplexity: PerplexityBot, plus Perplexity-User, which Perplexity says generally ignores robots.txt since a person requested that fetch. Google: Googlebot for Search, and Google-Extended for Gemini apps and Vertex AI grounding. These lists change, so read each provider's documentation rather than any blog post, this one included.
No. Google-Extended governs Google's Gemini apps and its Vertex AI grounding. Google's documentation states it does not affect inclusion in Google Search, and AI Overviews sit inside Search rather than beside it. What governs how much of your page can be shown in a Search result, including an AI Overview, are the ordinary snippet controls: nosnippet, max-snippet, and data-nosnippet. Blocking Google-Extended is a real decision about Gemini and Vertex. It does not control AI Overviews.
There are three common causes and they need different fixes. Your crawlers may be blocked, and now you know to check the right agent rather than just the training one. Your pages may be readable but shaped like a brochure, in which case the engine has nothing quotable to lift and the fix is writing pages that answer one question each. Or your business may be readable and clear but absent from the third-party sources the engines already trust, which is the slowest of the three and the only one that usually needs outside help.
It matters because it is the only description of the image a machine ever receives, unless the surrounding text happens to describe it. On my mom's shop, roughly 990 product images carried no alt text at all, which meant a large part of the catalog was invisible to anything that could not see pictures. Alt text is a translation layer. Write what the photo shows in one plain sentence, using only facts the listing itself supports, and skip the keyword stuffing.
No, and I would not start with one. Everything in this check uses your browser, a free AI chat account, and Google's Rich Results Test. Paid AI-visibility tools mostly automate the asking and log it over time, which is genuinely useful once you have a baseline and a reason to track it, and premature when you have not yet checked whether your crawlers are blocked. Do the free version first.
Quarterly is a reasonable rhythm for a small store, and after any theme change, platform migration, or privacy-plugin install, because those are the events that reintroduce a crawler block without anyone noticing. The one part worth running monthly is the first step, asking the engines your customer questions, since the answers shift as the models update. Log what you asked and what came back, or you will not be able to tell drift from noise.
This check is becoming one of the channel's first videos, run live on a real store: subscribe at youtube.com/@itsryanlenk to catch it. The receipts from running the full method on my mom's shop, 143 AI citations between June 10 and July 23, 2026 from Microsoft Copilot and its partner surfaces, on $0 in ads, are in the full case study article. If you want the rest of the site in the order to use it, start here.