askchatstudio Run the free checkFree check

Answer engine optimization · Original data

Your robots.txt is fine. Your firewall is the problem.

Every guide about getting your business into AI answers starts the same way. Go look at your robots.txt. Make sure you're not blocking GPTBot.

So I checked, and not on my own site. I took fifty-five actual local businesses, the kind ChatGPT names when someone asks for a recommendation in their city. I grabbed the domains straight from answers the assistant gave, so this isn't some random website list. These are businesses the model already knows about.

Out of forty-nine that responded, exactly one had robots.txt configured to block AI search crawlers. That is one site out of forty-nine. That's two percent.

If the advice were right, I should have found a pile of them. I found almost nothing.

Then I checked something else, and that number was nine.

The door isn't in the file

Here's the thing about robots.txt. It's a polite request. It lives at the root of your website and tells robots where they're welcome, and well-behaved crawlers read it and follow along. OpenAI's search crawler respects it. So does Anthropic's, and Perplexity's.

But a request only works if the robot actually reaches your site.

So instead of reading the file, I knocked on the front door. I requested the homepage of all fifty-five sites the way a crawler would, then did it again identifying as OAI-SearchBot, the crawler that feeds ChatGPT's answers.

Nine of the fifty-five wouldn't open. That's sixteen percent, which is eight times the robots.txt number.

What came back wasn't a page. Six sites returned 403 Forbidden, meaning the server saw the request and decided not to serve it. Two returned 429, too many requests, from a crawler that had made exactly one. One returned 418, a joke status code that bot-protection tools use to say no. And one didn't answer at all.

Only one of those nine also had anything in robots.txt. The other eight had a perfectly friendly file sitting there, inviting crawlers in, on a server that slams the door before anyone reads it.

Nobody chose this

I want to be clear about what's happening here, because it isn't neglect.

Every one of those sites is running some kind of bot protection. Came with the hosting plan, or a security plugin someone installed after a scraping incident, or a firewall rule a developer added years ago that nobody's looked at since. The setting was made against content scrapers and spam traffic, and it works fine for that.

AI assistants just ended up on the same side of the wall.

The owner has no idea. There's no warning anywhere. Your site loads perfectly in your browser, because your browser runs JavaScript and passes whatever challenge the protection throws at it. A crawler never does any of that. It asks for the page, gets a 403, and leaves.

And here's what makes it genuinely hard to catch. The assistant never tells anyone. It doesn't say "I tried to check this business and couldn't get in." It just writes its answer from whatever it could reach, and that is everybody else's pages.

They still show up, and that's the trap

This is what I didn't expect. All nine of those businesses appear in ChatGPT answers anyway. That's how they landed in my list.

Think about what that means. The assistant recommends them, describes what they do, tells a stranger whether they're worth calling, and every word comes from directories, review sites, and articles other people wrote. The business itself contributed nothing, because it couldn't.

So they're visible and voiceless at the same time. Every competitor with an open site gets to put their own words in front of a potential client. These nine get described by whoever wrote about them last, maybe three years ago, maybe wrong.

One of them is a business I track for a client. Their site returns 403 to the crawler while they pay someone every month to improve their pages. Those pages have never been read by the thing they're trying to reach.

How to check yours in two minutes

You don't need a tool for this and you don't need me. Open a terminal and run this, with your own domain:

curl -I -A "OAI-SearchBot/1.0" https://yoursite.com/

You're looking at the first line of what comes back. If it says 200, you're fine and the door is open. If it says 403, 418, or 429, the crawler is being turned away. If nothing comes back at all, that's the same problem wearing a different hat.

Do it for a few of your important pages too, since the homepage alone can mislead you. I've seen protection configured per-path, where the front page is open and the service pages aren't.

If you find a 403, the fix isn't on your website. It's in whatever sits in front of it. In Cloudflare you're looking for a bot-fight setting and a rule to allow verified bots. In WordPress it's usually a security plugin with a bot list. Your host's support team will know the specific switch if you tell them you need to allow AI search crawlers, and it takes them about five minutes.

While you're there, know which robots you're actually talking about, because they aren't the same thing. OAI-SearchBot, ChatGPT-User, Claude-SearchBot and PerplexityBot are the ones that fetch pages to answer questions. Those are the ones you want in. GPTBot, ClaudeBot and CCBot collect pages for training models. Blocking those is a completely reasonable decision and it doesn't cost you a single recommendation.

Most advice out there lumps them together and tells you to open everything. You don't have to. You can keep your content out of training sets and still be read by the assistant your client is asking.

The uncomfortable bit

I keep running into this pattern and it's the same shape every time. The work people do on AI visibility is mostly the visible stuff, things like writing pages, fixing headings, adding schema markup. The things that actually stop you are boring and invisible, and they sit in a settings panel nobody remembers opening.

Two percent of these businesses made a choice about AI crawlers. Sixteen percent had the choice made for them by a default.

Go check which group you're in. Takes two minutes and it's the only thing on your list tonight that might be worth eight times what you expected.

Method: 55 domains of independent local businesses named by ChatGPT in recommendation answers, pulled from our own tracking data. robots.txt fetched and parsed for the four AI search crawlers; homepages requested with a browser user agent and again as OAI-SearchBot. Raw results and the survey script are in our repo. Checked 21 September 2026.