---
title: 在 Next.js 中分别配置 AI 搜索与训练爬虫
description: 保留 ChatGPT 与 Claude 的搜索发现能力，同时分别决定是否允许训练爬虫和用户触发的 AI 访问。
canonical_url: https://nextaiready.com/zh/docs/guides/robots-txt
url: https://nextaiready.com/zh/docs/guides/robots-txt
last_updated: 2026-09-26
updated: 2026-09-26
author: next-ai-ready 团队
summary: 保留 ChatGPT 与 Claude 的搜索发现能力，同时分别决定是否允许训练爬虫和用户触发的 AI 访问。
topics: [nextjs, robots.txt, ai-search, ai-crawlers, chatgpt-search, claude-search]
---

# 在 Next.js 中分别配置 AI 搜索与训练爬虫

Next.js App Router 站点应使用 `app/robots.ts`，并分别决定 AI 搜索、模型训练和用户触发
读取的访问策略。如果目标是获得 AI 搜索可见性、但不希望内容参与模型训练，就允许搜索和
用户访问，同时阻止训练爬虫。

```ts
// app/robots.ts
import type { MetadataRoute } from "next"
import { aiRobots } from "next-ai-ready/robots"

export default function robots(): MetadataRoute.Robots {
  return aiRobots(
    { name: "Acme", baseUrl: "https://acme.com" },
    {
      aiBots: {
        search: "allow",
        user: "allow",
        training: "disallow",
        other: "allow",
      },
      sitemap: true,
    },
  )
}
```

这是一个有限选项的访问决策，不是 SEO 分数。它只说明哪些类型的 Agent 可以抓取站点，
不保证抓取、索引、排名、检索或引用。

## 按用途选择策略

不要把所有 AI user agent 当作同一种爬虫。主要服务商已经为不同用途提供了独立身份。

| 用途      | 代表性 Agent                                          | 对可见性的影响                               |
| ------- | -------------------------------------------------- | ------------------------------------- |
| AI 搜索   | `OAI-SearchBot`、`Claude-SearchBot`、`PerplexityBot` | 阻止后可能降低或失去在对应 AI 搜索中的发现机会。            |
| 模型开发    | `GPTBot`、`ClaudeBot`、`Google-Extended`             | 阻止后会退出各服务商说明的相关用途；这与普通 Google 搜索是两件事。 |
| 用户触发读取  | `ChatGPT-User`、`Claude-User`、`Perplexity-User`     | 阻止后，助手可能无法按用户要求读取页面。                  |
| 其他或混合用途 | 旧版和服务商特有 Agent                                     | 设置明确的 fallback，并在修改前核对服务商文档。          |

[OpenAI 官方文档](https://developers.openai.com/api/docs/bots)说明：`OAI-SearchBot` 用于
ChatGPT 搜索，`GPTBot` 用于模型开发，`ChatGPT-User` 由用户操作触发。OpenAI 还指出，
由于 `ChatGPT-User` 请求由用户发起，robots.txt 规则可能并不适用。

[Anthropic 官方文档](https://support.claude.com/zh-CN/articles/8896518-anthropic-%E6%98%AF%E5%90%A6%E4%BB%8E%E7%BD%91%E7%BB%9C%E7%88%AC%E5%8F%96%E6%95%B0%E6%8D%AE-%E7%BD%91%E7%AB%99%E6%89%80%E6%9C%89%E8%80%85%E5%A6%82%E4%BD%95%E9%98%BB%E6%AD%A2%E7%88%AC%E8%99%AB)
也采用三种身份：`Claude-SearchBot`、`ClaudeBot` 和 `Claude-User`。

`Google-Extended` 是 robots.txt 控制 token，并不是一个独立发送请求的 HTTP 爬虫身份。
[Google 明确说明](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers#google-extended)，
它不影响 Google 搜索收录，也不是 Google 搜索排名信号。

## 按用途配置静态文件

构建期生成的静态文件可以使用同一策略：

```js
// ai-ready.config.mjs
import { defineConfig } from "next-ai-ready"

export default defineConfig({
  site: {
    name: "Acme",
    baseUrl: "https://acme.com",
    description: "Acme 产品文档。",
  },
  robots: {
    aiBots: {
      search: "allow",
      user: "allow",
      training: "disallow",
      other: "allow",
    },
    sitemap: true,
  },
})
```

四种用途都使用有限选项：

| 字段         | 适用对象                  | 省略后的默认值                     |
| ---------- | --------------------- | --------------------------- |
| `search`   | AI 搜索索引爬虫             | 先读取 `default`，否则为 `"allow"` |
| `training` | 模型开发爬虫与控制 token       | 先读取 `default`，否则为 `"allow"` |
| `user`     | 用户触发的助手读取             | 先读取 `default`，否则为 `"allow"` |
| `other`    | 混合用途、旧版或未被独立说明的 Agent | 先读取 `default`，否则为 `"allow"` |
| `default`  | 未配置用途的 fallback       | `"allow"`                   |

需要严格白名单时，可以将默认值设为拒绝：

```js
robots: {
  aiBots: {
    default: "disallow",
    search: "allow",
  },
}
```

## 向后兼容的统一策略

原来的字符串写法仍然有效。只有当同一个决定确实适用于所有已知 AI bot 时才使用它：

```js
robots: {
  aiBots: "allow", // 或 "disallow"
}
```

未配置 `robots.aiBots` 时，next-ai-ready 保持原有默认行为，显式允许已知 AI bot。

## 静态文件还是 app/robots.ts

`next-ai-ready build` 默认写入 `public/robots.txt`。当策略只随部署变化时，这种方式适用于
任何托管平台。

对于 App Router 项目，Next.js 推荐使用 `app/robots.ts` 元数据文件约定。除非使用 Dynamic
API 或动态路由配置，否则它的返回结果会被缓存。`aiRobots()` 的结果兼容
`MetadataRoute.Robots`。

使用 `app/robots.ts` 时：

1. 设置 `emit: { robots: false }`，避免构建时再写一份 `public/robots.txt`。
2. 从 `app/robots.ts` 返回 `aiRobots(site, robotsConfig)`。
3. 只保留一套权威策略，避免静态规则和运行时规则互相冲突。

Doctor 能识别 `app/robots.ts` 与 `emit.robots: false`，不会误报缺少静态文件。

## llms.txt 注释不是标准指令

生成的静态文件包含指向 `/llms.txt` 和 `/llms-full.txt` 的注释。这些注释便于人类检查
机器入口，但 robots.txt 标准并不要求爬虫跟随它们。请保持真实入口可访问，并从站内页面
链接过去；不要把这些注释当作已经抓取或引用的证据。

## 验证生产响应

部署后检查真实文件：

```bash
curl -i https://example.com/robots.txt
```

不要只检查 HTTP 状态，还要检查实际分组：

```bash
curl -s https://example.com/robots.txt | grep -A1 -E \
  'OAI-SearchBot|GPTBot|Claude-SearchBot|ClaudeBot|Claude-User'
```

然后运行部署预检：

```bash
pnpm exec next-ai-ready audit https://example.com --version 3
```

Audit 验证的是技术访问能力。是否产生真实发现仍需结合 Search Console、服务端日志，以及
各服务商的引荐或引用数据来判断。

## 主要依据

- [Next.js robots.txt 元数据文件](https://nextjs.org/docs/app/api-reference/file-conventions/metadata/robots)
- [OpenAI 爬虫身份与控制方式](https://developers.openai.com/api/docs/bots)
- [Anthropic 爬虫身份与控制方式](https://support.claude.com/zh-CN/articles/8896518-anthropic-%E6%98%AF%E5%90%A6%E4%BB%8E%E7%BD%91%E7%BB%9C%E7%88%AC%E5%8F%96%E6%95%B0%E6%8D%AE-%E7%BD%91%E7%AB%99%E6%89%80%E6%9C%89%E8%80%85%E5%A6%82%E4%BD%95%E9%98%BB%E6%AD%A2%E7%88%AC%E8%99%AB)
- [Google-Extended 文档](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers#google-extended)
