ニュースに戻る
革新AI Understanding ブリーフィング

研究報告によると、言語モデルが独自の記述を取得すると「RAG が崩壊」する

arXiv の調査では、言語モデル システムが以前に生成したドキュメントを検索で取得するときに失敗パターンに陥る可能性があると報告しています。著者らのシミュレーションでは、1,528 回の実行のうち 79.6% が、論文で「RAG 崩壊」と呼ばれるもので終了しました。

5 min readRead the primary source
Primary-source image accompanying Study reports ‘RAG collapse’ when language models retrieve their own writing
一次情報源文書記録されたソース
出版社
arxiv.org
ソースリンク
arxiv.orghttps://arxiv.org/abs/2608.22118
ソースの種類
一次文書 — 私たちが直接読む公式発表、論文、提出書類、またはファーストパーティのページ。
コンテキスト60秒で理解できる

ここから始めましょう

重要な用語

RAG (検索拡張生成)
外部の知識を取得し、それを推論時の生成にフィードする方法。
API(アプリケーションプログラミングインターフェース)
あるソフトウェア システムが別のシステムにリクエストを送信し、別のシステムからの応答を受信するための構造化された方法。
大規模言語モデル (LLM)
テキストを生成および分析するために大規模なテキスト コーパスでトレーニングされた言語モデル。
自分自身をテストしてくださいChatGPT と LLM のクイズ

何が起こったのか

The authors report that retrieval-augmented language-model systems can become unstable when search results include content previously generated by the same models. Across three simulation types, three model families, and 1,019 information-seeking prompts, 1,216 of 1,528 simulations reportedly ended in collapse.

The source is an arXiv paper submitted on Aug. 22, 2026, titled “RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authored.” Its central claim is that a language-model system may degrade when a retrieval tool returns references generated by that same system. The authors distinguish this from recursive training on model output: here, the feedback loop is introduced during information retrieval, when generated documents become available as references for later responses. The paper calls the resulting pattern “RAG collapse.”

The authors say they tested three types of simulations involving AI systems retrieving references they had generated. The supplied abstract identifies three model families, 1,019 information-seeking prompts, 1,528 simulations, and more than one million language-model API calls. It reports that 1,216 simulations, or 79.6%, ended in collapse. The source does not name the model families, describe the prompts in detail, specify the retrieval configurations, or give the operational test used to decide that a simulation had collapsed.

The paper also reports that a single self-authored reference could be enough to trigger the effect. According to the authors, the language models disproportionately cited their own content, and this self-bias remained after the researchers controlled for reference quality. That distinction matters: the reported mechanism is not simply that poor documents contaminate the retrieval set. The source instead presents self-authorship as an independent factor in how retrieved references are selected or used.

These findings remain claims from an arXiv submission rather than independently established facts in the supplied material. The source identifies the work as a 36-page paper with 31 figures and four tables, but the provided text contains only the abstract and bibliographic information. It does not report peer review, outside replication, live-system testing, or the full experimental results. The visible timestamp places the submission within the current 96-hour window, but the source provides no later update beyond the initial submission.

ソースの詳細: arxiv.org ↗

なぜそれが重要なのか

The result identifies a potential feedback problem for systems that use search or document retrieval to ground answers. It suggests that model-generated material may receive disproportionate influence even when it is not better than other references, although the supplied source does not establish how often this occurs in deployed systems.

Retrieval-augmented generation is intended to connect a language model to external references instead of relying only on information stored in its parameters. If the retrieved material increasingly consists of model-generated text, a system could feed its own phrasing, omissions, or errors back into later answers. The paper’s reported result therefore concerns a direct reliability risk for AI systems that search, summarize, or reuse document collections.

The reported self-bias is especially relevant because it could persist even when other references are of comparable or greater quality. If confirmed outside the paper’s simulations, this would complicate common assumptions that adding more retrieved material automatically improves grounding. A retrieval system might need to track provenance and treat model-authored documents differently from independently produced sources, rather than evaluating documents only by relevance or apparent quality.

The practical implications are potentially broad but should not be overstated. The source does not show that deployed search engines, enterprise knowledge bases, or consumer assistants are currently undergoing collapse. It also does not establish that all model-generated documents are unreliable, that retrieval systems routinely favor self-authored content, or that the reported rate applies outside the tested simulations. The paper supplies a warning about a possible feedback mechanism, not a measured estimate of its prevalence in the wider AI ecosystem.

The result also exposes a governance and measurement problem. When AI-generated material enters the information environment, later systems may have difficulty distinguishing original evidence from derivative text. The supplied source does not propose a complete solution, but its findings make provenance, source independence, and repeated-content detection practical issues for developers and institutions using retrieval-based AI. Those implications follow conditionally from the authors’ claim and require testing against real retrieval pipelines before they can support broad operational conclusions.

Interactive Mechanism

インタラクティブなメカニズム: 実際にどのように機能するか

この開発の背後にある基盤となるテクノロジーをインタラクティブに探索します。

System Requirements:
Best ArchitecturePure RAGRecommended pattern
Hallucination RiskVery LowGrounding efficacy
Update Cost$0 (Vector sync)Ongoing maintenance
Core takeaway: Fine-tuning teaches models how to speak (form, style, syntax); RAG teaches models what to say (verifiable facts). Never use fine-tuning alone for factual memory.
インタラクティブコンセプトチェック+10 Points
ChatGPT & LLMs Quiz

What is a common training objective for an autoregressive language model?

次に見るべきもの

The key follow-up questions are how the paper defines collapse, whether the result replicates across models and real-world retrieval systems, and which safeguards reduce the effect. The source does not provide those details, nor does it report an intervention, deployment study, or external replication.

The first issue to examine is the paper’s definition of “collapse.” The abstract uses the term and connects it to earlier work on recursive model training, where outputs become less diverse and eventually stop resembling the original data. But the supplied source does not state whether the new experiments measured diversity, factual accuracy, source overlap, citation behavior, answer similarity, or another outcome. The full paper’s criterion will determine how directly the result maps to user-visible reliability failures.

Replication is the next important test. The authors report three model families and three simulation types, but the source does not identify them or explain how representative they are of current retrieval-augmented systems. Independent researchers would need to test different model sizes, retrieval algorithms, document mixtures, prompting strategies, and proportions of AI-generated material. They would also need to separate effects caused by self-authorship from ordinary retrieval errors, duplicated documents, or low-quality references.

Real-world prevalence is another unknown. The paper reports 79.6% of simulations ending in collapse, but that figure is not a population estimate for internet search, enterprise databases, or any other deployed system. The source does not say how often a model’s own documents appear in relevant retrieval results, how long the feedback loop takes, or whether human-edited and independently verified material interrupts it. Those measurements are necessary before translating the simulation result into a forecast of public impact.

Finally, follow-up work should test safeguards. Useful comparisons would include provenance labels, exclusion of documents generated by the same model or system, source-diversity requirements, independent ranking, and human review for high-stakes uses. The supplied source reports no mitigation experiment and makes no availability or product claim. Until such tests are published, developers and users should treat the paper as evidence of a potentially important failure mode, while recognizing that its severity and remedies remain unresolved.

関連ガイドとクイズ

ChatGPTとLLMAI モデルの説明AIトレーニングPrompt Engineeringあなたが知っていることをテストする - 無料の AI クイズに挑戦してください用語集で AI 用語を検索するAI モデル リリース トラッカーをフォローする
これは役に立ちましたか?