OpenAI announced Monday it will begin embedding an invisible watermark into text produced by ChatGPT and Codex for users in the European Union, according to a company blog post, marking the company's first deployment of text watermarking technology to comply with EU transparency mandates. The move responds to the EU AI Act's transparency requirements, which became enforceable on August 2 and obligate AI firms to label AI-generated material in ways that other systems can recognize. The watermark will deploy over the next several weeks to eligible ChatGPT and Codex subscribers across all subscription tiers within the EU, though OpenAI isn't making the feature mandatory worldwide at rollout.

The watermark operates by gently influencing the model's selection of words, creating a pattern invisible to human readers but identifiable by detection software, the company explained. Because the signal resides within the word choices themselves, it persists when text is copied and pasted to other platforms. Developers accessing OpenAI's API globally can activate the watermark for certain models starting today, though it remains disabled by default. OpenAI reported observing no significant degradation in model performance when the watermark was enabled, and emphasized the watermark doesn't reveal user identity. However, the company's own testing revealed vulnerabilities: swapping just 10% of words with synonyms caused detection accuracy to plummet from roughly 92% to 66%. Brief excerpts, mathematical responses, and translated material proved especially difficult to detect reliably.

"These limitations contribute to our decision to provide initial detector access only to approved researchers and expert organizations, who can help us evaluate reliability and responsible uses," the company stated. OpenAI also warned that an absent watermark "does not prove human authorship," since text might be too short, too extensively modified, or originate from a competitor's AI system. The company published a technical paper on its approach, dubbed textGrain, co-authored with scholars from the University of Pennsylvania and Yale, which demonstrates using a secret key to organize next-word predictions when completing sentences; hundreds of these subtle adjustments allow detectors to identify AI-generated content using only the text and the cryptographic key.

The announcement arrives two months after Anthropic said it would watermark Claude-generated text globally, a decision that sparked criticism from some Claude users who contended they had provided "the instructions, context, decisions" while Claude served merely as "the tool." OpenAI had previously developed a text watermarking system but delayed its release, partly due to worries that customers would migrate to competitors without watermarking requirements, The Wall Street Journal reported in 2024. According to the blog post, watermarks "can indicate that an OpenAI system generated or processed part of a passage, but not how much human judgment, editing, or creativity went into it." Anthropic, Google, Meta, Microsoft, and OpenAI have all pledged to adhere to the EU's code of practice governing AI-generated content. The company's EU-first strategy suggests regulatory pressure rather than competitive advantage is driving adoption, while the detector's fragility against simple synonym swaps raises questions about whether watermarking can remain effective as users learn evasion techniques.