Pew sampled the web, not a pile of recent blog posts
Pew drew 10,000 English-language pages from each of 49 Common Crawl snapshots between January 2021 and July 2026. That produced 490,000 pages spanning the period before and after ChatGPT's public release on November 30, 2022.
Common Crawl is a large archive of publicly accessible webpages. Paywalled and login-only material is likely underrepresented, so this is not every page a person might read online. It is still much closer to a broad web sample than a study built from one publisher, topic or social platform.
Pew stripped out the page text and ran it through editlens_Llama-3.2-3B, an open-weight detector made by Pangram. Pages scoring 0.2 or higher were counted as showing meaningful signs of AI authorship or editing. In the July 2026 random sample, roughly one in ten crossed that line.
The 35% headline needs its denominator
Only 10% to 15% of pages in each crawl had a publication date Pew could detect in the HTML. Pew used that subset to isolate pages published after ChatGPT arrived. In July 2026, 35% of those dated, post-ChatGPT pages showed signs of AI involvement.
That dated subset is not a random sample of the whole web, and Pew says so in the methodology. News stories, blog posts and commercial articles are more likely to carry clean publication metadata than an old help page, a product record or a page behind a login. The 35% figure tells us a lot about newer dated content. It does not mean a third of every page online was written by a bot.
The domain split adds another clue. Pew found signs of AI authorship on around 10% of .com pages in its 2026 samples, compared with 4.6% of .org pages and roughly 1% of .edu and .gov pages. The rise is concentrated where publishing more pages can support search traffic, sales or marketing. That is a plausible incentive story, not proof of why any particular company used AI.
A trend detector can work while a one-page verdict fails
Pew checked whether its result depended on the open model. It also ran 62,370 pages through Pangram's commercial model. The two systems agreed on 96% of cases, with a Cohen's kappa of 0.61. Their aggregate estimates followed the same upward direction.
The disagreements still matter. Pew found that the open model produced more apparent positives on pages from 2021 and 2022, before public AI writing tools were common. The report treats those early estimates as approximate and likely inflated by human writing that the detector misclassified.
That is the line worth keeping. Apply one method consistently across hundreds of thousands of pages and it can reveal a broad shift. Feed it one student's essay, one applicant's cover letter or one employee's report and the score is not a final answer about who wrote what. Pew explicitly warns that an individual classification should not be treated as a definitive verdict.
Yes, the writing tics are spreading
Pew also counted language features associated with model-written prose. Compared with its 2023 snapshot, em dashes appeared about twice as often. Oxford commas rose 63%. A basket of words including “delve,” “pivotal,” “crucial” and “tapestry” more than doubled. Negative parallelism—the tidy “it's not just X, it's Y” construction—nearly tripled, though it remained uncommon overall.
This is funny until normal writers start deleting punctuation to prove they are human. People used em dashes before ChatGPT. Editors used Oxford commas before ChatGPT. A phrase becoming common in generated text does not transfer ownership of that phrase to a machine.
The useful reading test is harder than spotting a verbal tic. Does the page make a claim specific enough to check? Do its links support that claim? Are dates, sample sizes and limits attached to the number? Does a named person or organization stand behind it? Bland machine prose often fails those checks. So does bland human prose.
Priya trusts the trend. Mina refuses the accusation.
Priya Rao sees a useful population-level result with two different denominators. She would keep 10% beside the random July sample and 35% beside the dated post-ChatGPT subset every time the study is cited. The larger number becomes misleading the second its narrower sample disappears.
Mina Torres draws the line at using this research against one person. If a school or employer suspects undisclosed AI use, a probabilistic detector should not become the case by itself. Let the person show drafts, notes, sources or version history and explain the work. A punctuation hunch is not due process.
Priya wants the trend reported cleanly. Mina wants the individual treated fairly. Those positions fit together. The web can be filling with AI-shaped writing while any single page remains uncertain.
What changes for people who publish and read the web
Writers do not need to perform humanity by making every sentence rougher. They do need to leave more evidence behind the page: sources that open, dates that can be checked, original reporting where they have it and clear ownership of mistakes. If AI materially shaped the work, a plain disclosure is more useful than forcing readers to play punctuation detective.
Readers can stop treating polish as proof. A smooth answer with no method can waste more time than a clumsy page with the right document attached. For consequential advice—money, health, school, law or a purchase—open the source and check when it was published. If the page cannot tell you where its confidence came from, close it.
Pew's study does not say the web is over. It says new writing is being produced under a different cost structure, with more of it carrying the same statistical fingerprints. The scarce part is no longer a clean paragraph. It is a page worth believing after you read it.