trustme.bro/r/…
✓ checked
trust me, bro:
here is the receipt.
the claim
AI-generated text poses contamination and quality threats to linguistic corpora.
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
2 sources for · 0 against

Peer-reviewed literature demonstrates that the proliferation of low-quality, AI-generated text contaminates open datasets and degrades information integrity within information ecosystems.

Evidence for · 2
2026 · cited by 0
Language models acquire useful linguistic and generative behavior from the statistical structure of their training data. As public and synthetic corpora increasingly contain duplicated, weakly filtered, low-information, and recursively generated text, it becomes important to measure how low-signal contamination affects model capability under controlled conditions. This paper presents an experimental study of capability degradation in small language models under controlled low-signal textual corruption. A clean corpus is used to construct fixed-size dataset variants with corruption ratios of 0%, 30%, 50%, 60%, and 80%. Corruption is generated through a formal taxonomy consisting of repetitive low-entropy text, semantically incoherent text, distractor-dominant text, recursive synthetic text, and shallow high-fluency filler. Small GPT-style decoder-only transformer models are trained from scratch under identical architecture, tokenizer, optimizer, training schedule, and evaluation settings across three random seeds per condition. Degradation is measured using clean-test perplexity, repetition behavior, lexical diversity, prompt adherence, bootstrap confidence intervals, benchmark-style multiple-choice probes, ablation analysis, and a Capability Retention Index. Results from 30 model-run rows show that mixed low-signal corruption increases mean clean-test perplexity from 1654.12 at 0% corruption to 2199.29 at 80% corruption, a 32.96% increase. CRI remains below the clean baseline for every corrupted mixed variant and reaches 0.6139 at 80% corruption, although it is not strictly monotonic under the current aggregate metric. Ablation results identify repetitive low-entropy and shallow fluent-filler corruption as the most harmful categories under the tested metrics.The term capability is used operationally in this study to denote retained language-modeling and generation behavior under fixed evaluation metrics, not broad human-like cognition or frontier-scale task competence.The study contributes a reproducible framework for data-quality degradation analysis while keeping claims limited to the tested small-model, three-seed setting.
See more details
The analysis

rails:sufficiency:supported:single_source:for=1+1p:against=0+0p | v55:sufficiency

More for · 1
2025 · cited by 0
Quarantining Cheap AI and Synthetic Exhaust: An Entropy-Based Imperative for Information IntegrityCivilization Physics — Entropy & Information Ecology Series This whitepaper argues that the accelerating flood of unregulated, low-integrity AI-generated content—“synthetic exhaust”—poses a civilizational-scale entropy threat to global information ecosystems. As cheap AI systems produce massive volumes of unverified text, code, and media, this synthetic content contaminates the very knowledge streams that future models, search engines, and human institutions depend on. Once mixed into open datasets, this entropy cannot be undone. Grounded in the Entropy Law (R) and the Information Inbreeding Hypothesis, the paper shows how recursive exposure to synthetic content creates a closed informational loop: rare signals disappear, noise amplifies, and models collapse into narrow, homogenized, low-fidelity outputs. Experiments already confirm that models trained even partially on AI-generated data rapidly degrade, losing diversity, accuracy, and world-model fidelity. As synthetic contamination spreads across the open internet, the entire AI pipeline risks drifting into an irreversible epistemic decline. Cheap AI systems accelerate this collapse because they combine:• zero Integrity (no grounding, no consistency mechanisms)• zero Presence (no human oversight, no judgment loops)• infinite output capacity (high-volume pollution at near-zero cost) The result is a global cascade of entropy: deg
Everything we examined (2)
This check searched the claim as stated. It did not run a separate search for evidence against it.
  1. Capability Degradation in Small Language Models Under Low-Signal Data Corruptionpeer-reviewedno side taken
  2. Quarantining Cheap AI and Synthetic Exhaust: An Entropy-Based Imperative for Information Integritypeer-reviewedno side taken
This receipt carries no identity, shared or not. Sharing publishes your connection to it, not your data.
Check your own claim
Challenge the receipt
trust me, bro: win the argument, pass the class, survive peer review.
This receipt is an automated verdict against our published method · not an opinion about any author or publication.
Terms · Privacy · How verdicts work · Dispute this receipt