hould You Block AI Bots? What Publishers Get Wrong

Should Publishers Block AI Crawlers? What Publishers Get Wrong

Blocking AI bots isn’t as simple as flipping a switch. See how publishers can protect content without losing search traffic in the process.

O que você vai encontrar neste artigo

If you are thinking about blocking ai crawlers from your website, hold that thought for just a second.

As we saw at the latest Refinery89’s Publisher At The Core discussion session, AI crawlers are becoming a big concern for publishers. They use publisher content to gather information to train AI models and generate answers, often without permission or compensation. At the same time, blocking these bots entirely could reduce your visibility in search results, potentially affecting traffic and revenue. So, is blocking every bot really the right call?

Here’s the catch: different bots do different jobs. Some index your site for search, some ground an AI answer, and some crawl purely to train a model. Treat them as one item, and you’ll likely make the wrong decision in either direction: blocking traffic you actually wanted or allowing the exact crawling you meant to stop.

Before you fall into the temptation of blocking them altogether, let’s take a step back and analyze the situation to decide whether it’s worth banning bots from the lands of your website or letting them roam free.

 

How did we get here?

 

In the past year, AI-powered search features have been reshaping discovery as we know it very quickly. Therefore, it has been taking a toll on publishers’ traffic and ad monetization performance. A February 2026 Ahrefs study found that when an AI Overview appears, click-through rates for top-ranking pages drop by 58% on average. That drop in clicks is also showing up in publisher revenue. Ozone data reported by Digiday showed that ad inventory across publisher sites in the US and the UK fell by as much as 40% year-over-year in Q2 2026.

 

Why you should think twice before blocking bots

 

Publishers are already feeling the pressure of the new era of search, which is putting their visibility, revenue, and integrity at risk. The obvious fix to protect their site’s content seems to be blocking bots. But this is not a one-size-fits-all solution since not all bots are used for training AI models or powering AI answers. Some search engine crawlers have the main purpose of helping index your page in search result, while others are used by demand partners to understand your ad inventory and maximize its value.

Not all bots serve the same purpose, and that distinction is what makes blocking risky. A single blanket block aimed at “AI” can end up catching everything at once, including the search crawler you still depend on for visibility.

 

There’s no industry consensus on what to do

 

Publishers are taking different approaches to this challenge, and they mostly depend on their size, growth strategies, and resources. During AdExchanger’s Programmatic AI 2026, Microsoft’s VP of publisher product told publishers to let bots in and pursue licensing deals instead. This statement raised the opposite argument that blocking all AI bots first and then only allowing the relevant ones is what gives publishers the leverage to negotiate licensing deals. That strategy, however, tends to work best for publishers whose content is valuable enough that AI companies are willing to pay for access.

 

Why robots.txt isn’t the wall you think it is

 

Blocking bots is not the unbreakable barrier it seems to be. As we explained in our blog on how to use robots.txt as part of your website strategy, this is just a request and it’s up to the bots to honor it or not. There’s not 100% effective enforcement mechanism behind a disallow rule. Kind of like placing a “Do not disturb” sign rather than a locked door.

This detail is important because it changes the idea of what blocking accomplishes. A publisher can believe they’ve shut a crawler out entirely, while it keeps crawling anyway, with no way to verify compliance. Compliance varies significantly by crawler, so a rule that works against one bot may do nothing against another wearing a different name. Let’s see some ways to enforce the access decisions for your site.

 

How can publishers protect their content from AI bots?

 

New tools are being developed and updated to help publishers some control over which bots get access their content. However, none of them (at least not on their own, for now), will solve the whole issue. Here are some you can keep in mind:

 

Google Search Console’s AI opt-out toggle

Activating this toggle means you will not receive traffic or impressions from Google’s generative AI features (AI Overviews and AI Mode). However, this control is not a ranking signal for search results meaning you can still appear on the traditional search engine result pages.

 

Robots.txt

A robots.txt file tells search engine and AI crawlers which parts of a website they can or cannot access. It helps publishers control how their content is crawled and a decent bot will respect these instructions. However, as we mentioned earlier, it relies on compliant crawlers and is not a security measure that will protect your content from being scraped. You can read more about this in our blog.

 

Cloudfare

Cloudflare is an infrastructure provider used to improve website performance and security. It has been developing tools to help control bot access more efficiently. Realizing that their current solutions were too blunt due to the different types of bots that exists, the company decided to group them into three separate crawler categories:

  • Buscar: indexes content to answer questions later. It is still expected to send referral traffic
  • Agent: acts in real time on a user’s behalf (ChatGPT-User, Claude/Gemini browsing).
  • Training: absorbs content permanently to improve a model.

Starting in September 2026, new Cloudflare sites will default to blocking Training and Agent bots. This only applies to pages with ads. Search bots will stay allowed, but crawlers like Googlebot fall into more than one category (Search and Training in this case), meaning they will be blocked. However, publishers can opt out of this. You can read the full announcement here.

 

Best Practices for Managing AI Crawlers and Protecting Publisher Revenue

 

Audit your crawler logs

Before writing any allow or disallow rules, check your server logs to see what bots are actually hitting your site and what their main job is (indexing, retrieval, or training) before deciding what to block.

 

Separate helpful bots from unwanted ones

Tools that now separate search, agent, and training traffic let you block what doesn’t add value. They so this without cutting off the search visibility you still depend on.

 

Do not block what your pages need to load

Keep your CSS, JavaScript, and images crawlable, because search engines rely on them to understand how your page is built.

 

Use blocking as leverage, not a goal

A block only becomes a negotiating position if it’s paired with a licensing proposal. Otherwise, it’s traffic given up for nothing.

 

Continuously review your decision

Crawler behavior, blocking norms, and technologies are moving so fast. A policy, best practice, or even a solution set today may be outdated within a few months. Therefore, it’s recommended to stay informed about evolving market and keep track of your own site’s performance to adjust your strategy accordingly.

 

Check in with your AdTech partner

When in doubt, talk to your AdTech partner to make sure you are not blocking any crawlers that influence your ad monetization. At Refinery89, we can help you make sure all your demand partners are able to monetize your site effectively and flag any issues we spot.

 

Blocking AI Crawlers: Finding the Right Balance

 

So, to block or not to block? That really is the question, and the answer depends on your site’s capabilities. AI bots serve different purposes. Some scrape content, others train AI models, and some help your website gain visibility in search. Opting for licensing deals or only allowing valuable bots to visit your site are two valid methods that you can adapt to your business goals if you have the means to do so correctly.

We highly recommend talking to your AdTech partner before doing any work on your site. This way you will ensure your traffic and ad monetization are not impacted in the process. They’ll guide you toward finding the best solution to protect your content while also protecting your ad revenue. If you want more insights into preparing your site to endure the new era of AI search, reach out to our experts at Refinery89.

 

Escrito por
Compartilhar:
WhatsApp
E-mail
Facebook
X
LinkedIn
Reddit
Receba as últimas novidades
Assine Nossa Newsletter Semanal

Mantenha-se atualizado com as últimas notícias, treinamentos e webinars

Mais Popular
Navegue por Categoria!