Skip to content
Approvalens

Reading room · 5 min read

The noai Meta Tag: Does It Stop AI Training? What Actually Works

Where noai and noimageai came from, which crawlers document support for them (none of the big ones), and the opt-outs OpenAI, Google, Apple and Bing honour.

By the Approvalens team

Fixes these report findings

  • noai directives

<meta name="robots" content="noai, noimageai"> is a signal DeviantArt introduced in 2022 to say "don't use this content to train AI". It isn't part of any standard, Google lists it nowhere among the robots rules it supports, and none of the crawler documentation from OpenAI, Anthropic, Perplexity, Google, Apple or Amazon mentions it. Leaving it on your pages does no harm, but if you want an opt-out the big AI companies document, use their robots.txt tokens (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot) and, for answer engines, the snippet controls they do support.

Where noai came from

In November 2022 DeviantArt added two directives to pages whose artists opted out of AI datasets: noai and noimageai. TechCrunch reported at the time that the tags were appended to the HTML page of each artwork, and that DeviantArt's updated terms required third parties training on DeviantArt content to exclude pages carrying them. DeviantArt also encouraged other creator platforms to adopt similar protections.

So noai was a platform policy backed by one site's terms of service. It wasn't agreed with the companies that run AI crawlers. If you find it on your own pages and didn't add it, look at your theme, your plugins and any header rules on your server or CDN.

Who documents support for it

Operator What their documentation says to use Mentions noai?
OpenAI GPTBot in robots.txt for training; OAI-SearchBot for ChatGPT search (OpenAI crawlers) No
Anthropic ClaudeBot in robots.txt for training (Anthropic crawlers) No
Google Google-Extended in robots.txt for Gemini training and grounding; nosnippet and related rules for Search, including AI Overviews (robots meta tag) No
Apple Applebot-Extended in robots.txt for model training; nosnippet for AI-generated answers in Siri and Search (About Applebot) No
Amazon Robots meta noarchive means "do not use the page for model training" (Amazonbot) No
Microsoft Bing nocache and noarchive control use in Bing's chat answers and in training (Bing Webmaster Blog) No
Common Crawl CCBot in robots.txt (CCBot) No

Google's robots meta documentation lists the values it supports (all, noindex, nofollow, none, nosnippet, indexifembedded, max-snippet, max-image-preview, max-video-preview, notranslate, noimageindex, unavailable_after) and publishes the same list in a machine-readable file. noai isn't on it.

"Doesn't mention" isn't the same as "ignores". A company could honour the tag without documenting it. But you can't verify that, and an opt-out you can't verify isn't one you should rely on.

What to use instead

To keep content out of AI training

Add robots.txt groups for each training token:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

These are the tokens each company documents for training. None of them affects search: OpenAI, Google and Apple each state that their training token is separate from search inclusion. The allow AI crawlers in robots.txt guide has the full table.

To limit what answer engines quote

  • Google Search, including AI Overviews and AI Mode: nosnippet "will also prevent the content from being used as a direct input for AI Overviews and AI Mode". Use data-nosnippet on specific elements, or max-snippet to limit length.
  • Apple: nosnippet means Apple "will not use data tagged nosnippet as additional context" for AI-generated output in its products.
  • Bing: per Microsoft's 2023 announcement, nocache limits chat answers to URL, title and snippet; noarchive keeps content out of chat answers and training. Pages still appear in Bing search.
  • Amazon: noarchive means "do not use the page for model training".

Snippet rules have a cost: nosnippet also removes the text snippet from your normal Google search result. Use it on the pages or elements you actually want to protect, not sitewide.

Should you remove noai if you have it?

Keep it if you want; it signals intent and costs nothing. Remove it if it was added by a plugin you no longer understand, or if you actually want AI tools to use your content and the tag is sending the opposite message. Either way, don't let it replace the robots.txt rules above.

Check what your pages actually send. The tag can be in the HTML:

<meta name="robots" content="noai, noimageai">

or in an HTTP header, which you won't see in View Source:

curl -sI https://example.com/ | grep -i x-robots-tag

If you find noindex next to noai in the same tag, that's a bigger problem: noindex removes pages from Google Search. The noindex by mistake guide shows how to find where it's coming from.

noai and AdSense

Unrelated. AdSense uses its own crawler and doesn't read noai. A stray noindex beside it is worth fixing for search, and a free Approvalens scan flags noindex on pages you want reviewed.

Check your pages

The free AI crawler checker looks for noai, noimageai and similar directives in both the robots meta tag and the X-Robots-Tag header, and shows your robots.txt rules for each AI bot next to them, so you can see whether your opt-out uses a signal the crawlers document. Its page requests use simulated AI user agents from our server, so a firewall that checks IP ranges may respond differently to the real crawlers.

FAQ

Does noai stop ChatGPT from reading my site?

OpenAI's crawler documentation doesn't mention it. To opt out of training use GPTBot in robots.txt; to stay out of ChatGPT search use OAI-SearchBot.

Is noimageai any different?

It's the image-specific variant from the same DeviantArt policy. Same status: not in Google's supported list and not mentioned by the AI crawlers' documentation.

Does noai hurt my Google rankings?

Google lists the values it supports and noai isn't one, so there's nothing in Google's documentation suggesting it affects Search.

What about other new opt-out signals?

Several proposals exist, including Cloudflare's Content-signal line in robots.txt (managed robots.txt). Use them if you like, but check each AI company's own documentation before counting on them.

Spotted something out of date or wrong? Tell us and we'll correct it.

Read this guide in Turkish →

Check it on your own site. Free, no sign-up.

Free tools for this

Free scan

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and every problem found · no sign-up

All guides →