A customer asks ChatGPT who to call. Does your business get named?
Search behavior has split. Some customers still type into Google and scan blue links. A growing share ask an assistant a question and act on whatever name comes back. That second group never sees your ad, your map pin, or your five-star reviews — they see whatever the model decided to say out loud. This is a plain accounting of which crawlers feed that answer, what each one is actually named, and what a business with no marketing department can do about it.
What decides whether an AI names your business instead of a competitor's?
Three things, in order: whether the assistant's crawler can reach and read your site at all, whether your basic facts (name, address, phone, hours, services) are stated in plain unambiguous text somewhere it can find, and whether other sources — directories, review sites, local press — say the same facts back. Design and traffic rank do not matter to this process the way they matter to Google's blue links.
Every major answer engine works the same way at a mechanical level: a program requests your page, reads the text, and either uses it to answer a live question or folds it into a training run. If that program is blocked — by a robots.txt rule, by a site that requires JavaScript to render any visible content, by a page with no text describing what the business actually does — nothing downstream happens. You cannot be cited from a page that was never read.
Assuming access, the next filter is clarity. These systems are extracting facts, not vibes. A page that states 'Licensed electrician serving Port St. Lucie, Fort Pierce, and Vero Beach, open Monday through Saturday, emergency calls after hours' gives a model something to quote. A hero image with the same information baked into a graphic gives it nothing — no crawler reads text off a JPEG.
The last filter is corroboration. Answer engines lean on aggregate signal the same way people do: if your Google Business Profile, your own site, an Angi listing, and a local news mention all agree you're an electrical contractor at the same address with the same phone number, that agreement is itself a ranking input. One glossy site with no outside confirmation reads as less trustworthy than five plain, matching listings.
Which crawlers actually feed ChatGPT, and what does each one do?
OpenAI runs four separately named crawlers with different jobs: GPTBot gathers training data, OAI-SearchBot powers ChatGPT's search feature, ChatGPT-User fetches a page live when a user's question triggers it, and OAI-AdsBot checks advertiser landing pages. Each can be allowed or blocked independently in robots.txt.
OpenAI's published bot documentation is specific enough to act on. GPTBot is the training crawler — blocking it opts your content out of future foundation-model training but does not affect whether ChatGPT can still surface your page in a live search answer. OAI-SearchBot is the one that matters for the search feature inside ChatGPT: a site that disallows it can still appear as a plain navigation link, but it won't be pulled into a generated answer. ChatGPT-User is different again — it fires in the moment, when a live conversation asks the assistant to go check a specific page, and it is not a background crawl.
Practically, almost no small business should block any of these. The failure mode owners actually hit isn't an aggressive block rule — it's a robots.txt inherited from a template, or a site builder that disallows all bots by default because nobody ever opened the file. Check yours. It sits at yourdomain.com/robots.txt and is worth reading once, in full, rather than trusting whatever the platform shipped.
What about Perplexity, Claude, and Google's AI Overviews specifically?
Perplexity runs PerplexityBot for indexing and Perplexity-User for live answer fetches — its own documentation states neither is used to train foundation models. Anthropic runs ClaudeBot, Claude-User, and Claude-SearchBot for the same three purposes. Google's AI Overviews use standard Googlebot indexing plus Google-Extended, a separate toggle that controls AI training and grounding use without affecting normal Search ranking.
The naming is consistent across vendors once you see the pattern: one crawler for background indexing, one for live per-query fetches, and — where the vendor also trains models — a third, separate one just for training data. That separation exists specifically so a site can allow being cited in a live answer while opting out of training, or vice versa. Read each vendor's bot page directly rather than guessing; user-agent strings and behavior have changed before and will again.
Google is the case worth being precise about, because it's the one people misconfigure hardest. Google-Extended has nothing to do with whether Googlebot can crawl and index you for regular Search — it only governs whether crawled content feeds Gemini training and certain grounding uses. Blocking it does not hurt your ranking. Google's own AI-features documentation is also blunt that there is no special schema or markup requirement to appear in AI Overviews — the standard SEO fundamentals (crawlable, indexed, factually clear) are what it evaluates.
| Vendor | Crawler | Job |
|---|---|---|
| OpenAI | GPTBot | Foundation-model training |
| OpenAI | OAI-SearchBot | ChatGPT search-feature indexing |
| OpenAI | ChatGPT-User | Live fetch during a user's conversation |
| OpenAI | OAI-AdsBot | Advertiser landing-page checks |
| Perplexity | PerplexityBot | Search indexing (not model training) |
| Perplexity | Perplexity-User | Live fetch during a user's question |
| Anthropic | ClaudeBot | Model training and development |
| Anthropic | Claude-User | Live fetch during a user's query |
| Anthropic | Claude-SearchBot | Search-result indexing |
| Googlebot | Search indexing (Search, Images, Discover) | |
| Google-Extended | AI training and grounding use, separate from ranking |
Does adding schema markup actually get you cited?
Structured data helps a machine parse your facts correctly, but it is not a guaranteed switch. Google states directly that no special schema is required to appear in AI Overviews — the standard is crawlable, indexed, and factually clear content. Google also fully retired classic FAQPage rich results in 2026, so 'add FAQ schema for a Google snippet' is now outdated advice.
This is the point where a lot of AI-visibility advice gets hypey, so it's worth being blunt: schema.org markup — LocalBusiness type with name, address, telephone, and hours filled in — makes your facts unambiguous to any program reading the page, human-built or machine. That's a real, durable benefit. What it is not is a lever that forces a citation. Google's own current AI-features guidance says explicitly that no new markup, AI text files, or special schema is needed to appear in AI Overviews; the standard SEO basics are what get evaluated.
The FAQPage case is the clearest example of advice going stale. Through 2023, Google restricted the classic FAQ rich-result snippet to government and health sites. By mid-2026 the feature was deprecated and the documentation removed entirely — it no longer shows in Google Search results for anyone. If a vendor is still selling 'FAQ schema gets you a rich snippet' as a 2026 tactic, that claim is out of date. Question-and-answer formatting still helps because it matches how people phrase queries to an assistant — that's a content-structure benefit, not a schema-markup one.
The honest framing: structured data is hygiene, not magic. It removes ambiguity for anything parsing your page. It does not substitute for the page actually containing the answer in plain language.
What should a local business actually do this week?
Open robots.txt and read it. Confirm nothing is blocking GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, Claude-User, Claude-SearchBot, or Googlebot unless you have a specific reason to. Most accidental blocks come from a platform default, not a decision anyone made.
Write your core facts in plain text somewhere a crawler can read without executing JavaScript: business name, service area, phone, hours, and what you actually do, in sentences — not only inside a hero graphic or a contact form. Add LocalBusiness schema with the same facts. It won't force a citation, but it removes any ambiguity for whatever does the reading.
Match your NAP — name, address, phone — exactly across your site, Google Business Profile, and every directory listing you control. Mismatched suite numbers or a swapped phone number are the single most common reason corroboration fails silently.
Write a handful of pages that answer a real question in the first two sentences, the way you'd answer it out loud to a customer on the phone. Assistants quote clear, self-contained answers far more readily than marketing copy that takes three paragraphs to say what you do.
Get named somewhere you don't control — a local news mention, an industry directory, a chamber listing. Third-party corroboration is the input you cannot manufacture solely on your own site, and it's the one that separates a business an AI trusts from one it merely found.
Questions we get asked on this
- Most local businesses shouldn't. Blocking GPTBot, ClaudeBot, or Google-Extended opts you out of AI training use, but it also removes you from the pool of sources those systems can cite by name. For a business trying to get found, being readable is the goal — blocking is a decision for content creators actively worried about training use, not for a contractor trying to get named when someone asks an assistant who to call.
- No, and the distinction matters. GPTBot is OpenAI's training crawler, run on its own schedule independent of any single user. ChatGPT-User fires in real time, only when someone in an active conversation asks the assistant to check a specific page. You can allow one and block the other in robots.txt; they are controlled separately.
- Not directly — those are different companies with their own crawlers and their own indexes. Ranking well in Google Search doesn't automatically transfer trust to Perplexity or ChatGPT's search feature. What does transfer is the underlying reason you rank: clear, factual, consistently stated information, which every one of these systems independently rewards.
- The classic Google FAQ rich-result feature was deprecated in 2026, so FAQ schema no longer earns a special snippet in Google Search. Writing content in a clear question-and-answer structure is still worth doing because it matches how people phrase questions to an assistant — just don't expect the markup itself to produce a visual result in Google anymore.
- Check your robots.txt file directly at yourdomain.com/robots.txt for any Disallow rule under the crawler names covered above. Separately, if your site relies entirely on JavaScript to render visible text, some crawlers may retrieve an empty or near-empty page. A quick way to sanity-check this is viewing your page's source HTML directly — if your core facts aren't in that raw source, a text-only crawler may not see them either.
- Consistency of your name, address, and phone number across every place they appear — your site, Google Business Profile, and any directory listing. It's the cheapest fix, it's within your control without touching code, and it's the corroboration signal every one of these systems checks before trusting a fact stated on your own site alone.
Sources
- OpenAI — bot and crawler documentation (GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot)
- Perplexity — crawler documentation (PerplexityBot, Perplexity-User)
- Anthropic / Claude — crawler support article (ClaudeBot, Claude-User, Claude-SearchBot)
- Google — common crawlers reference (Googlebot, Google-Extended)
- Google — AI features documentation (no special markup required for AI Overviews)
- Google — LocalBusiness structured data guidance
- Google — FAQPage structured data (deprecated 2026)
- Schema.org — LocalBusiness type reference
Last updated 2026-07-20.
Find out what the crawlers currently see on your site.
Send us your domain and we'll tell you what's actually blocked, missing, or inconsistent — no guessing, checked against the same documentation cited above.
Keep exploring: AI receptionist · Local SEO fundamentals · Map pack ranking · Getting found online
