Pew study finds more than a third of new web pages show signs of AI writing
The Pew Research Center examined nearly half a million English-language web pages. Since ChatGPT launched in late 2022, the proportion of machine-written text online has risen sharply.
Researchers used the Common Crawl web archive and checked samples with the Open Pangram detection tool. In a July 2026 sample, about 10 percent of all pages examined showed clear signs of AI authorship.
The figures change when looking only at pages published after ChatGPT’s release. More than a third of those newer pages show signs of AI authorship. The trend began with ChatGPT in late 2022, and the share of likely AI-generated web content has climbed steadily since then.
Commercial websites are roughly ten times more likely to contain AI-written text than pages from schools or government agencies. About one in ten pages with a .com domain shows signs of AI authorship. .org domains sit at 4.6 percent. .edu and .gov domains come in at only about 1 percent each.
“Delve,” em dashes, and Oxford commas are booming
Pew’s analysis found several language patterns that have become much more common on the web since 2023. Em dashes now show up about twice as often as they did in 2023. Oxford comma usage has jumped 63 percent.
Certain AI-favorite words like delve, interplay, testament, pivotal, landscape, tapestry, bolstered, crucial, meticulous, and vibrant have more than doubled in frequency. Negative parallelisms following the it’s not just X, it’s Y pattern have nearly tripled, though they remain rare in absolute numbers. A separate study looking at corporate PR documents found that this particular phrase quadrupled since 2022.
A study by Imperial College London, the Internet Archive, and Stanford University from April 2026 reached a similar conclusion. Researchers found roughly 35 percent of all newly published websites were fully or partly AI-generated. They also found 33 percent higher semantic similarity between AI texts and a much more positive tone overall. The team cautioned that public perception of negative effects often goes well beyond what the data actually supports.
What counts as “AI text” remains fuzzy
There is a problem with this and similar studies, and with the public debate too. Nobody agrees on what AI text even means. The spectrum runs from fully automated content to human drafts polished with AI to texts where a model only stepped in for a few sentences.
Open Pangram and similar models can, in my experience from testing hundreds of my own texts, only make a very rough call on whether a human or a machine likely wrote something. They can’t reliably tell you how much AI was involved or at what stage, and they still misfire regularly. Yet these are very different ways of working.
Public debate around AI text is growing more polarized, as the recent discussion about Anthropic‘s planned watermark for Claude output made clear. Using AI tools already carries a stigma that cuts both ways, something workplace studies have documented as well. Neither camp leaves much room for the messy reality of how people actually write with these tools. And since AI adoption isn’t slowing down, figuring out that middle ground is going to matter more and more.




