ニュースに戻る
革新AI Understanding ブリーフィング

Mugglehead は、ChatGPT 以降に公開されたウェブページの 35% に AI 関連の書き込みサインがあると報告しています

Mugglehead は、ピュー研究所の分析で、ChatGPT の開始後に公開された英語 Web ページの約 35% で AI の書き込みに関連する重大な兆候が見つかったと報告し、検出結果は AI がページ全体を生成したことを証明するものではないと警告しています。

5 min readRead the linked source
Source-provided image accompanying Mugglehead reports AI-associated writing signs on 35% of webpages published after ChatGPT
出典参照記録されたソース
出版社
mugglehead.com
ソースリンク
mugglehead.comhttps://mugglehead.com/artificial-intelligence-signs-appear-on-35-of-webpages-published-since-chatgpt/
ソースの種類
リンクされたソース — プライマリ ソースのステータスが確立されていません。
コンテキスト60秒で理解できる

ここから始めましょう

重要な用語

合成データ
機密トレーニング データを強化、シミュレート、または保護するために使用される人工的に生成されたデータ。
生成AI
テキスト、画像、オーディオ、ビデオ、コードなどの新しいコンテンツを生成する AI システム。
データセット
トレーニング、検証、テストに使用される構造化サンプルまたは非構造化サンプルのコレクション。
自分自身をテストしてくださいAIとは何ですか?クイズ

何が起こったのか

Mugglehead Investment Magazine reports that Pew Research Center analyzed nearly 490,000 webpages collected from the Common Crawl archive and found significant signs associated with AI authorship on about 35% of pages published after ChatGPT launched in November 2022. Across the full period studied, from January 2021 through July 2026, AI-associated patterns appeared on about one in 10 English-language webpages. The source does not independently verify Pew’s underlying analysis or .

Mugglehead reports that the Pew Research Center examined nearly 490,000 webpages collected from the Common Crawl web archive. The pages were published between January 2021 and July 2026, allowing researchers to compare writing patterns before and after ChatGPT launched in November 2022. According to the report, AI-associated signs appeared on about one in 10 English-language webpages across the broader , but the rate rose to approximately 35% among pages published after ChatGPT’s launch. These figures describe the share of pages flagged by a detection system; they do not establish that AI generated all or even most of the text on each page.

The report says researchers used Open Pangram, an AI-detection model developed by Pangram Labs. Rather than searching for individual words or phrases, the software looks for statistical patterns across large amounts of writing. Mugglehead reports that the results varied sharply by domain: .com pages showed signs of AI writing at about 10 times the rate of .edu and .gov pages, whose rates were near 1%. About 4.6% of .org pages showed significant signs of AI involvement. The four domain categories had similarly low recorded rates in 2021, according to the report.

Mugglehead also describes changes in writing patterns associated with . Em dashes appeared about twice as often as before 2023, while use of the Oxford comma increased by 63%. Words such as “delve,” “interplay” and “testament” more than doubled in frequency, and negative parallelism—constructions such as “not just X, but Y”—nearly tripled. The report says that pattern remains uncommon despite its rapid increase. It also reports that AI-associated writing on .com pages rose from roughly 1% in January 2021 to 9.35% by January 2026.

ソースの詳細: mugglehead.com ↗

なぜそれが重要なのか

The findings suggest that machine-assisted writing may be entering the web at a scale relevant to search, publishing and future AI training data. Mugglehead reports that commercial .com pages showed these signs at roughly 10 times the rate of .edu and .gov pages. If machine-generated material is repeatedly included in future training sets, researchers warn that errors, bias and lost information could be amplified, although the study cannot establish that every flagged page was written by AI.

The reported increase matters because the open web is both a publishing environment and a source of material for later information systems. Mugglehead says Common Crawl estimates that its archive supplies 70% to 90% of the training-data tokens used by nearly all major large language models. If that estimate and the reported detection results are representative, a growing share of web text may be produced or altered with AI assistance before being collected for future model development. The practical concern is not simply that readers encounter synthetic prose; it is that synthetic prose may become part of the background material used to build later systems.

The source connects this possibility to model collapse, a term used for degradation that can occur when AI systems are repeatedly trained on machine-generated material. Mugglehead reports that researchers have warned repeated training on can reduce the information preserved by a model and can reinforce existing errors, biases and unfairness. University of Toronto computer engineering professor Nicolas Papernot compared the process to repeatedly photocopying a photocopy. That comparison conveys the reported risk, but the source does not provide a quantified estimate of how much web-based synthetic content would be required to cause degradation or whether current commercial training pipelines are already experiencing it.

The domain differences also point to an uneven distribution of AI-assisted publishing. Mugglehead reports that commercial sites had far higher detection rates than academic and government sites, which the article links partly to differences in editorial review, institutional approval, publishing speed and scale. Commercial domains include news outlets, marketing operations and automated content businesses, according to the report. Those categories are broad, however, and the source does not show how results differ among journalism, advertising, product pages or other types of .com content. The findings therefore support concern about where AI-assisted writing may be concentrated, not a blanket conclusion about the quality or origin of commercial webpages.

Interactive Mechanism

インタラクティブなメカニズム: 実際にどのように機能するか

この開発の背後にある基盤となるテクノロジーをインタラクティブに探索します。

System Requirements:
Best ArchitecturePure RAGRecommended pattern
Hallucination RiskVery LowGrounding efficacy
Update Cost$0 (Vector sync)Ongoing maintenance
Core takeaway: Fine-tuning teaches models how to speak (form, style, syntax); RAG teaches models what to say (verifiable facts). Never use fine-tuning alone for factual memory.
インタラクティブコンセプトチェック+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

次に見るべきもの

The key issue is whether AI-detection systems can become reliable enough to distinguish generated text from human writing assisted by AI. Mugglehead reports that the Open Pangram model detects statistical patterns rather than particular words, and that detection can produce false classifications. Future reporting should examine Pew’s methods, the detector’s validation, how results vary by topic and language, and whether Common Crawl or model developers introduce safeguards against training repeatedly on synthetic text.

The largest uncertainty is what a positive detection result means. Mugglehead explicitly reports that AI-detection software can incorrectly classify both human-written and machine-generated text. A flagged page may contain writing edited or assisted by AI rather than text generated entirely by a model. The source also does not independently confirm the Pew analysis, identify the detector’s measured false-positive and false-negative rates, or explain how multilingual pages, copied text, revisions and search-engine content were handled. Those methodological details will determine how much confidence readers should place in the 35% figure.

Detector performance may also change as writing conventions and model outputs evolve. The report says Anthropic has reportedly worked on model-level text fingerprinting for Claude outputs, which could identify generated writing without relying entirely on linguistic patterns. Mugglehead does not provide evidence that such fingerprinting is deployed, effective across systems, or available for the pages in Pew’s sample. Follow-up work should compare statistical detectors with provenance tools, assess human-AI collaboration separately from fully generated text, and test whether common stylistic signals remain useful after writers and models adapt.

Future coverage should also track how web archives and model developers respond. Questions include whether Common Crawl can label or filter synthetic material, whether publishers disclose AI assistance, and whether model-training pipelines distinguish human-authored, AI-assisted and machine-generated text. The source offers no evidence that any of these safeguards are currently in place. It also does not establish that the reported increase caused measurable harm to an existing model. The most defensible near-term conclusion is narrower: Mugglehead reports a substantial rise in pages exhibiting statistical signals associated with AI writing, while the scale, causes and downstream effects remain incompletely known.

関連ガイドとクイズ

AIとは何ですか?AI モデルの説明AIトレーニングAI倫理あなたが知っていることをテストする - 無料の AI クイズに挑戦してください用語集で AI 用語を検索するAI モデル リリース トラッカーをフォローする
これは役に立ちましたか?