Claude 如何标记 AI 生成的内容
How Claude marks AI-generated content

原始链接: https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content

Anthropic 已承诺遵守《欧盟人工智能法案》的实践准则,通过为人工智能生成的内容实施机器可读的透明度措施。自 2026 年 8 月 2 日起,所有新版 Claude 模型都将嵌入文本水印和符合 C2PA 标准的来源签名元数据。这些标记将适用于全球所有 Claude 产品及云合作伙伴。 该公司也正致力于为旧版模型追溯性地添加这些功能。为确保透明度,Anthropic 计划提供工具供用户和第三方检测这些标记,以标识内容是由 Claude 处理的。 不过,Anthropic 强调这些标记具有局限性:它们并非来源的决定性证明,因为内容在经过深度编辑、重新保存或总结后,可能会导致标记被移除或遮盖。此外,没有标记也并不代表内容一定是由人类创作的。Anthropic 将发布更多技术文档,以协助开发人员履行其在《欧盟人工智能法案》下的透明度义务。

近期的一场 Hacker News 讨论聚焦于 Anthropic 为 Claude 生成的内容添加标识的策略。为符合即将实施的欧盟法规,Anthropic 正在将“难以察觉的水印”直接植入生成文本中。与简单的元数据不同,这些水印被编织在内容本身,使其在复制粘贴和轻微编辑后依然存在。 评论者们对该方法的性质及其有效性展开了讨论。一些用户推测,水印可能涉及用词或结构上的细微模式,并指出 Claude 独特的写作风格本身就已经带有一种“AI 感”。另一些人则提出了更简单的方法,例如使用独特的 Unicode 字符,但他们也指出这些字符很容易被清理工具去除。 尽管一些人对该功能如何应用于生成的代码表示关注,但普遍的共识是,虽然此类措施确实能更轻松地识别 AI 输出,但它们并非万无一失。精通技术用户很可能会找到规避这些标记的方法,但这一举措在解决有关 AI 透明度及生成内容潜在滥用的问题上,仍是重要的一步。
相关文章

原文

Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, as a provider of both generative AI models and generative AI systems. This article describes how we’re planning to put those commitments into practice, how marking works, and what its limitations are. We’ll update this article and publish more detailed technical guidance as it becomes available.

Anthropic’s commitments under the EU AI Act’s Code of Practice on Transparency of AI-Generated Content

What our marking commitments mean for Claude:

  • New models will mark AI-generated content from day one. Claude models launched in the EU on or after August 2, 2026 will support machine-readable marking at launch. Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported.

  • Marking works everywhere you use Claude. Marks will apply to output from supported Claude models across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and wherever Claude is offered, worldwide. Some platforms or features may not support certain marking types.

  • We'll help you detect Claude's marks. We'll support users and other third parties to detect Claude’s marks, as the Code requires, and we’ll share details in forthcoming documentation.

  • Existing models are in progress. The law includes a transition period for Anthropic models launched before August 2, 2026, and we’re working to add marking support for those models as well.

More details about our marking plans are below.

Machine-readable marks in Claude-generated content

As AI-generated content becomes commonplace, greater transparency and signals about where content comes from can give people useful context about the information they consume. To support transparency and comply with our legal obligations, Anthropic is working to include machine-readable marks in content that Claude generates.

What’s covered

  • Models. Claude models launched on or after August 2, 2026 support marking at launch. We’re also working to add marking support to Claude models released before that date, and we’ll update this article as that becomes available.

  • Products. Claude markings cover output from supported models everywhere you use Claude, including Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag. Embedded watermarks will apply to all generated text. Provenance metadata will apply where Claude supports processing files.

  • Cloud partners. Embedded watermarks will apply when supported Claude models are accessed through AWS, Google Cloud, or Microsoft Foundry. Signed provenance metadata may not be supported on every platform, depending on the features each platform offers.

  • Regions. Marking will apply to output from supported models wherever Claude is offered, worldwide.

How Claude marks content

Claude uses two complementary techniques to mark content generated and processed by Claude: (1) watermarks embedded in text, and (2) signed provenance metadata attached to files.

1. Embedded watermarks in text

When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.

Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from.

2. Signed provenance metadata

When Claude generates a supported file type, such as a .svg, .png, or .jpg, it will attach signed provenance metadata. This metadata follows the Coalition for Content Provenance and Authenticity (C2PA) open standard, which is used across the industry to record information about content provenance. If a signed metadata label is present, it signals that a file was processed by Claude and lets you detect whether the file has been tampered with.

Detecting Claude’s marks

We’re also working to enable users and other third parties to detect Claude’s embedded watermarks and provenance metadata. Detection checks whether a piece of text or a file carries a supported Claude mark. If a supported mark is found, it indicates that the content may have been processed by Claude.

We’ll share details on detection mechanisms in forthcoming technical documentation.

Limitations

Machine-readable marks provide important signals about content, but it’s worth understanding their limitations across all content types.

  • A detected mark provides a signal that content was processed by Claude, but is not fully conclusive. Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content. For example:

    • Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source;

    • The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.

  • Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed. Content generated by Claude may not carry a detectable mark if, for example:

    • It was generated by a model released before marking was supported;

    • The text has been heavily edited, paraphrased, translated, or mixed into other writing;

    • The passage is very short, leaving too little text for a reliable signal;

    • A file’s metadata was stripped through format conversion, re-saving, screenshots, or other means;

    • It was produced through a platform, feature, or file type where a particular marking type wasn’t supported.

If you build with Claude

If you deploy Claude in your own product, you should independently assess what Article 50 requires of your products and services. Consistent with our commitments under the EU Code, our goal is to support you in meeting your own transparency obligations, and we'll share technical guidance on our marking and detection approach as it becomes available.

联系我们 contact @ memedata.com