---
title: Configure robots.txt for AI search and training crawlers in Next.js
description: Keep ChatGPT and Claude search discovery available while making a separate robots.txt decision for training and user-triggered AI access.
canonical_url: https://nextaiready.com/en/docs/guides/robots-txt
url: https://nextaiready.com/en/docs/guides/robots-txt
last_updated: 2026-09-26
updated: 2026-09-26
author: next-ai-ready team
summary: Keep ChatGPT and Claude search discovery available while making a separate robots.txt decision for training and user-triggered AI access.
topics: [nextjs, robots.txt, ai-search, ai-crawlers, chatgpt-search, claude-search]
---

# Configure robots.txt for AI search and training crawlers in Next.js

Use `app/robots.ts` for a Next.js App Router site and make separate decisions for AI search,
model training, and user-triggered retrieval. If your goal is AI visibility without opting into
training, allow search and user crawlers while disallowing training crawlers.

```ts
// app/robots.ts
import type { MetadataRoute } from "next"
import { aiRobots } from "next-ai-ready/robots"

export default function robots(): MetadataRoute.Robots {
  return aiRobots(
    { name: "Acme", baseUrl: "https://acme.com" },
    {
      aiBots: {
        search: "allow",
        user: "allow",
        training: "disallow",
        other: "allow",
      },
      sitemap: true,
    },
  )
}
```

This is a bounded policy decision, not an SEO score. It says which classes of agents may fetch the
site; it does not guarantee crawling, indexing, ranking, retrieval, or citation.

## Choose a policy by purpose

Do not treat every AI user agent as the same kind of crawler. The major providers expose separate
identities for different jobs.

| Purpose                  | Representative agents                                | Visibility consequence                                                                                        |
| ------------------------ | ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| AI search                | `OAI-SearchBot`, `Claude-SearchBot`, `PerplexityBot` | Blocking can reduce or prevent discovery in the provider's search experience.                                 |
| Model development        | `GPTBot`, `ClaudeBot`, `Google-Extended`             | Blocking opts content out of the uses described by each provider; it is separate from ordinary Google Search. |
| User-triggered retrieval | `ChatGPT-User`, `Claude-User`, `Perplexity-User`     | Blocking can prevent an assistant from fetching the page for a user's request.                                |
| Other or mixed           | Legacy and provider-specific agents                  | Use an explicit fallback and review provider documentation before changing it.                                |

[OpenAI documents](https://developers.openai.com/api/docs/bots) that `OAI-SearchBot` is for ChatGPT
search, `GPTBot` is for model development, and `ChatGPT-User` is user initiated. OpenAI also notes
that robots.txt rules may not apply to `ChatGPT-User` because the request is initiated by a person.

[Anthropic documents](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler)
the same three-way split: `Claude-SearchBot`, `ClaudeBot`, and `Claude-User`.

`Google-Extended` is a robots.txt control token rather than a separate HTTP crawler identity.
[Google states](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers#google-extended)
that it has no effect on Google Search inclusion and is not a Google Search ranking signal.

## Purpose-based configuration

Configure the same policy for the build-time static file:

```js
// ai-ready.config.mjs
import { defineConfig } from "next-ai-ready"

export default defineConfig({
  site: {
    name: "Acme",
    baseUrl: "https://acme.com",
    description: "Acme product documentation.",
  },
  robots: {
    aiBots: {
      search: "allow",
      user: "allow",
      training: "disallow",
      other: "allow",
    },
    sitemap: true,
  },
})
```

The four keys are finite choices:

| Key        | Applies to                                            | Default when omitted      |
| ---------- | ----------------------------------------------------- | ------------------------- |
| `search`   | AI search indexing crawlers                           | `default`, then `"allow"` |
| `training` | Model-development crawlers and control tokens         | `default`, then `"allow"` |
| `user`     | User-triggered assistant fetches                      | `default`, then `"allow"` |
| `other`    | Mixed, legacy, or not independently documented agents | `default`, then `"allow"` |
| `default`  | Fallback for an omitted purpose                       | `"allow"`                 |

Use `default: "disallow"` when the site requires an explicit allowlist:

```js
robots: {
  aiBots: {
    default: "disallow",
    search: "allow",
  },
}
```

## Backward-compatible all-or-nothing policy

The original string form remains supported. Use it only when the same decision genuinely applies
to every known AI bot:

```js
robots: {
  aiBots: "allow", // or "disallow"
}
```

With no `robots.aiBots` setting, next-ai-ready keeps its existing default and explicitly allows
known AI bots.

## Static file or app/robots.ts

`next-ai-ready build` writes `public/robots.txt` by default. This works on any host and is useful
when the policy changes only with a deployment.

For an App Router project, Next.js recommends the `app/robots.ts` metadata-file convention. Its
return value is cached unless it uses a Dynamic API or dynamic route configuration. The result from
`aiRobots()` is compatible with `MetadataRoute.Robots`.

When using `app/robots.ts`:

1. Set `emit: { robots: false }` so the build does not also write `public/robots.txt`.
2. Return `aiRobots(site, robotsConfig)` from `app/robots.ts`.
3. Keep only one authoritative policy so static and runtime rules cannot conflict.

Doctor recognizes `app/robots.ts` with `emit.robots: false` and does not warn about a missing static
file.

## llms.txt comments are not directives

The generated static file includes comments pointing to `/llms.txt` and `/llms-full.txt`. These
comments make the machine-facing entrypoints easier for people to inspect, but no robots.txt
standard requires a crawler to follow them. Keep the real discovery paths available and linked from
the site; do not report the comments as proof of crawling or citation.

## Verify the deployed response

Inspect the production file after deployment:

```bash
curl -i https://example.com/robots.txt
```

Check the actual groups, not only the HTTP status:

```bash
curl -s https://example.com/robots.txt | grep -A1 -E \
  'OAI-SearchBot|GPTBot|Claude-SearchBot|ClaudeBot|Claude-User'
```

Then run the deployed preflight:

```bash
pnpm exec next-ai-ready audit https://example.com --version 3
```

The audit verifies technical access. Search Console, server logs, and provider-specific referral or
citation data are still required to measure whether the policy produces real discovery.

## Primary references

- [Next.js robots.txt metadata file](https://nextjs.org/docs/app/api-reference/file-conventions/metadata/robots)
- [OpenAI crawler identities and controls](https://developers.openai.com/api/docs/bots)
- [Anthropic crawler identities and controls](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler)
- [Google-Extended documentation](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers#google-extended)
