Is the GPTBot crawling your content yet?

code API

BANNED - Failed to resolve a dispute
Joined
Oct 20, 2018
Messages
6,376
Reaction score
3,514
If you’ve published content online recently, there’s a good chance GPTBot has already crawled it.

GPTBot is OpenAI’s web crawler that collects publicly available data to help train and fine-tune its large language models (LLMs), like the one powering ChatGPT. That means it’s helping artificial intelligence (AI) learn from your blog posts, product pages, help docs, and more. But should you let it?

Source: the man with the bald head, I know most of you guys don't like him :D
originally posted on Neil Patel's blog: https://neilpatel.com/blog/what-is-gpt-bot/

Edit: I have noticed some of content is getting crawled by chatgpt on some sites and it is interesting to see how AI integrates into SEO. I hope this has some positive effects to printing money overtime. Only time will tell.
 
Last edited:
Yeah, its crawling it's own content, what goes around comes around. Funny how karma works now even with ai.
 
Yeah, noticed GPTBot crawling some sites too. Curious if it’ll help SEO or traffic long-term. Guess we’ll see.
 
Chat GPT has already suggested some videos of my YT channels in their answers. Is there any method to make use of this situation (to earn money) ?
 
Yeah, GPTbot is crawling a lot of public sites now. It's interesting to see how AI and SEO are starting to mix. But only time will tell if it really helps boost rankings or revenue.
 
You can check by reviewing your server logs for requests from GPTBot’s user agent. If it’s allowed by your robots.txt, it might already be crawling your site.
 
Yeah, its crawling it's own content, what goes around comes around. Funny how karma works now even with ai.

I'd be surprised if OpenAI doesn't filter out the obvious AI-generated posts with a detector before training on it
 
ChatGPT sometimes recommend answers from your site too [if relevant to the query of course] with a link bank to your site as the source and that is the interesting source.

I only hope they won’t be as greedy as Google that they start to bastardize answers with ads. I can understand product recommendations but not for informational and educational content.

For now, let’s continue to hope it serves the right purpose ORGANICALLY.
 
I let it crawl. Blocking GPTBot means missing out on potential AI-driven traffic sources that might emerge. Plus, if your content gets referenced in AI responses, that's essentially free exposure.
The real question: Will AI citations become a ranking factor? Google's already testing AI overviews, so content that trains these models might get preference.
You can block it in robots.txt if you're worried, but honestly the horse has already left the barn on AI training data.
 
Back
Top