间谍标记,而非水印
Spymarks, Not Watermarks

原始链接: https://brand.io/article/spymarks/

“水印”(watermark)一词正被重塑为“间谍印记”(spymark),以突显其日益严重的隐私威胁。与用于证明所有权或真实性的传统可见水印不同,“间谍印记”是嵌入图像、音频、视频和文本中的隐形信号,用于追踪用户的身份、位置和行为。 谷歌(通过 SynthID)和 OpenAI 等科技巨头正越来越多地利用这些强大且难以察觉的算法,将数据库标识符直接编码进媒体文件中。由于这些信号被集成在文件的基础数据中(如像素模式、频域或特定的措辞选择),即使在常规元数据(如 EXIF 或 ID3 标签)被删除或文件经过编辑后,它们依然存在。 作者认为,尽管各大公司将此技术包装为识别 AI 生成内容的手段,但它实际上是一种强大的监控工具。通过在每一次数字交互中编码个人标识符,这些平台构建了一个系统,使用户在不知情且未同意的情况下被追踪,并使其社交网络被映射出来。采用“间谍印记”这一术语,旨在揭露该技术的本质:一种对个人隐私和言论自由构成重大风险的秘密监控机制。

这篇 Hacker News 帖子讨论了“间谍标记”(Spymarks)的概念,这是一种用于追踪单个内容副本的数字水印替代方案。 讨论探讨了这些追踪方法的局限性。一位用户质疑了间谍标记的有效性,认为它们很容易被绕过。对于图片,他们指出模拟复制或添加新图层可以掩盖原始标记;对于文本,他们建议只需重新录入内容或使用人工智能进行改写,即可有效地清除任何嵌入的标识符。 另一位评论者澄清说,“水印”一词在此语境下早已被使用,并指出其在数字媒体下载和电影预览片中用于追踪分发的历史应用。讨论还涉及了内容保护机制与用户规避手段之间持续的猫鼠游戏。
相关文章

原文

There’s a sneaky new evolution of the “watermark”, let’s call it — the spymark.

A watermark is a visible mark embedded in a physical or digital medium to verify authenticity or assert ownership.

A spymark is a hidden signal that makes your work traceable without your knowledge or consent.

Spymarking is sneaking onto the internet

Google SynthID is a spymark that embeds secret hidden signals “imperceptible to humans” (Google’s own words) into images, audio, text, and video. This signal can encode database identifiers that map to your identity. Your user records, full name, IP addresses, date of birth, physical addresses, political party affiliation, and more.

Google’s SynthID-Image paper reports that its SynthID-O variant can encode a 136-bit payload in a 512x512-pixel image. That is enough room for a 64-bit database identifier, with 72 bits left for error correction.

SynthID was not the first spymark system designed, and it’s hardly the only one under active development. OpenAI and many other tech companies are developing these systems at scale. These companies claim spymarking can help identify AI-generated content, yet they’ve gone beyond simple watermarking and built in robust tracking mechanisms. Social media, content production tools, and smartphones may soon find themselves filled with spymarking algorithms that sneak these signals into everything you publish.

For example, images can be invisibly altered in their frequency domain to carry tracking information such as database IDs linked to users:

A smaller version of Mochi’s photo contains a real toy watermark: ID 173, encoded through subtle pixel changes. Compare the original and marked image, animate the amplified differences, and decode the ID from the PNG. The associated author and timestamp are fictional.

Why give it a new name?

“Watermark” has become a catchall for historical marks of authenticity, banknote security features, copyright overlays, and hidden tracking signals in our media. That last usage obscures the privacy risk that some forms of modern “watermarks” have become.

Most people tire of the discourse on privacy. It’s complex, repetitive, and seems irrelevant to the typical day-to-day routine of the average citizen. We can’t keep explaining the technicalities behind statistical tracking tools each and every time we need to communicate these concepts to people and expect them to pay attention.

“To speak the name is to control the thing.” — Ursula K. Le Guin, The Rule of Names

With a simple change of terminology we can put the privacy concern up front, cementing this issue in the conversation forever:

  • Spy- — clandestine surveillance, involuntary disclosure
  • -mark — embedded signal

Now we get to disambiguate friendly watermarks from concerning new technology while leveraging the familiar etymology. It collapses a technical communication salient and gains us a lot of ground with just one new word. It’s clear, coherent, and straight to the point.

No more wasting energy establishing the basic facts. The privacy concern is built into the word.

“Spymark” is a privacy Rumpelstiltskin.

More spymark examples

Here are a few more forms of spymarks so you can better familiarize yourself.

Audio spymarks operate using principles similar to image spymarks, and they are typically inaudible. Some methods make minute changes to the audio waveform in the time domain; others modify features in the frequency domain, and some combine both approaches. These methods can encode data, including identifiers linked to personal information:

Audio watermark spectrograms

Compare spectrograms of the same LJ Speech excerpt unwatermarked and with Timbre, AudioSeal, WavMark, FSVC, Patchwork, or Norm-Space. The authors’ original plots are linked below.

Spectrograms from Wen et al. (2025), SoK: How Robust is Audio Watermarking in Generative AI models? · Original samples. Plots retain the authors’ original scales.

These schemes are engineered to be robust. They can often survive compression or re-encoding.

These tools are already proliferating, and in fact they predated generative AI and the push to watermark generative media. The open source spymarking tool audiowmark originated in 2018 and can hide 128-bit payloads in audio and protect them with a secret AES key, preventing users without the key from decoding them.

You can even encode personal information invisibly into text! SynthID steers word choices to create a detectable statistical pattern that can encode a tracking payload:

An illustrative encoding: eight word choices represent eight bits. The binary value 10101101 gives database ID 173, which can point to a record containing an author and timestamp. These are fictional examples, not a SynthID decoder.

Watermarks are benign, spymarks are not

Watermarks remain easy for the user to spot and are typically not nefarious.

Sometimes watermarks deter counterfeiting:

A $20 bill with a circular close-up of its faint portrait watermark.

Sometimes watermarks claim ownership:

Dorothea Lange’s Migrant Mother with a visible Getty Images watermark across the photograph.
Dorothea Lange, Migrant Mother (1936) — Getty sells some public domain images.

And sometimes they’re just annoying:

A lightly deep-fried Willy Wonka meme with TOP TEXT and BOTTOM TEXT captions, a small imgflip.com watermark in the bottom-left corner, www.9gag.com at the top right, and a tilted ifunny.co stamp across the middle.

But these watermarks don’t track you.

Printer tracking dots, on the other hand, are an early example of spymarks, dating back to the 1980s:

Diagram of printer tracking dots annotated to show encoded time, date, and printer serial number.
Image: EFF — Robert Lee, Seth Schoen, Patrick Murphy, Joel Alwen, and Andrew “bunnie” Huang

We need to be clear that spymarks are a form of metadata you have limited knowledge of and control over, and that their purpose is entirely antagonistic to you.

Metadata such as the EXIF tags in photos and the ID3 tags in MP3 files are standardized, well-documented fields. While EXIF can expose sensitive data such as GPS coordinates, these standardized fields can be inspected, edited, and removed from files you control.

Inspect editable EXIF and ID3 metadata in two tabs. EXIF shows fictional camera-maker, model, and author tags in a JPEG metadata segment; ID3 shows an MP3’s title, artist, and album. Hover or tap the bytes for explanations. These samples contain metadata only.

You can strip standard metadata tags out of your files. Unfortunately, a spymark signal embedded in pixels, audio, or word choices is invisible to you and can remain in your files even after you edit them.

Furthermore, whereas tags are helpful for maintaining information such as song titles or a photo’s exposure settings, spymarks encode user-tracking identifiers that are entirely opaque and useless to you. These signals can survive metadata removal and some edits, allowing marked copies to remain traceable as they circulate.

Keeping our files free of spying

It’s hard to imagine the future where we can’t even trust our own files. And the sad thing is that this has already started.

Spymarks are certainly not great for whistleblowers or anyone who doesn’t want to be persecuted for their words or affiliations. No matter where you stand on whatever issues, spymarks can be used against you and those you care about.

Imagine a future where every device is attested and every social media post carries an account-linked spymark. Where every file or post served to you for sharing has spymarks embedded within. Combined with account records and observations of where copies appear, these identifiers can help reconstruct how content spreads. In this world, it would be dangerously easy to hunt down anyone from a JPEG or a tweet. Or track the web of humans through which “dangerous ideas” flow.

That future lies halfway between now and 1984. So let’s stay off that timeline, shall we?

Call a spymark what it is. A spy tool used to spy on you and everyone you interact with.

— Brandon Thomas

Usage

spy·mark /ˈspaɪmɑːrk/ noun

plural spymarks

A hidden signal embedded in media to trace its origin, tools, or distribution history without the user’s meaningful control.

“The spymarks tied each copy to a different recipient.”

spymark verb, transitive

spymarked; spymarking; spymarks

To embed a spymark in a file or piece of media.

“The social network spymarks every uploaded image.”

spymarking noun

The practice of embedding spymarks.

“The platform introduced spymarking without telling its users.”

spymarked adjective

Containing a spymark.

“Removing the file’s metadata still left the image spymarked.”

联系我们 contact @ memedata.com