AI-generated writing can be distinguished from human writing
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
9 sources for · 0 against
Peer-reviewed literature and empirical studies establish that AI-generated writing can be distinguished from human writing using detection tools, stylometric markers, and linguistic cue analyses.
Natural language processing (NLP) has been studied in computing for decades. Recent technological advancements have led to the development of sophisticated artificial intelligence (AI) models, such as Chat Generative Pre-trained Transformer (ChatGPT). These models can perform a range of language tasks and generate human-like responses, which offers exciting prospects for academic efficiency. This manuscript aims at (i) exploring the potential benefits and threats of ChatGPT and other NLP technologies in academic writing and research publications; (ii) highlights the ethical considerations involved in using these tools, and (iii) consider the impact they may have on the authenticity and credibility of academic work. This study involved a literature review of relevant scholarly articles published in peer-reviewed journals indexed in Scopus as quartile 1. The search used keywords such as “ChatGPT,” “AI-generated text,” “academic writing,” and “natural language processing.” The analysis was carried out using a quasi-qualitative approach, which involved reading and critically evaluating the sources and identifying relevant data to support the research questions. The study found that ChatGPT and other NLP technologies have the potential to enhance academic writing and research efficiency. However, their use also raises concerns about the impact on the authenticity and credibility of academic work. The study highlights the need for comprehensive discussions on the potential use, threats, and limitations of these tools, emphasizing the importance of ethical and academic principles, with human intelligence and critical thinking at the forefront of the research process. This study highlights the need for comprehensive debates and ethical considerations involved in their use. The study also recommends that academics exercise caution when using these tools and ensure transparency in their use, emphasizing the importance of human intelligence and critical thinking in academic work.
Since ChatGPT has emerged as a major AIGC model, providing high-quality responses across a wide range of applications (including software development and maintenance), it has attracted much interest from many individuals. ChatGPT has great promise, but there are serious problems that might arise from its misuse, especially in the realms of education and public safety. Several AIGC detectors are available, and they have all been tested on genuine text. However, more study is needed to see how effective they are for multi-domain ChatGPT material. This study aims to fill this need by creating a multi-domain dataset for testing the state-of-the-art APIs and tools for detecting artificially generated information used by universities and other research institutions. A large dataset consisting of articles, abstracts, stories, news, and product reviews was created for this study. The second step is to use the newly created dataset to put six tools through their paces. Six different artificial intelligence (AI) text identification systems, including "GPTkit," "GPTZero," "Originality," "Sapling," "Writer," and "Zylalab," have accuracy rates between 55.29 and 97.0%. Although all the tools fared well in the evaluations, originality was particularly effective across the board.
Abstract
This study investigates techniques for detecting machine-generated text, a critical task in the era of advanced language models. We compare two approaches: a hand-crafted feature-based method and a deep learning method using RoBERTa. Experiments were conducted on diverse datasets, including the Human ChatGPT Comparison Corpus (HC3) and GPT-2 outputs. The hand-crafted approach achieved 94% F1 score on HC3 but struggled with cross-dataset generalization. In contrast, the RoBERTa-based method demonstrated superior performance and adaptability, achieving 98% F1 score on HC3 and 97.68% on GPT-2. Our findings underscore the need for adaptive detection methods as language models evolve. This research contributes to the development of robust techniques for identifying AI-generated content, addressing critical challenges in AI ethics and responsible technology use.
Large Language Models (LLMs) are used widely for tasks involving text generation such as dialogue summarization and creative writing. The generated text often appears unnatural, and this text can easily be distinguished from natural language. In this paper, we leverage the capabilities of Reinforcement Learning to fine-tune LLMs so as to produce text that resembles human language. We have applied the Proximal Policy Optimization algorithm to fine tune a FLAN-T5 LLM for a dialogue summarization task.
Large language models (LLMs) are now routine writing tools across various domains, intensifying questions about when text should be treated as human-authored, artificial intelligence (AI)-generated, or collaboratively produced. This rapid review aims to identify cue families reported in empirical studies as distinguishing AI from human-authored text and to assess how stable these cues are across genres/tasks, text lengths, and revision conditions. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA) guidelines, we searched four online databases for peer-reviewed empirical articles (1 January 2022–1 January 2026). After deduplication and screening, 40 studies were included. Evidence converged on five cue families: surface, discourse/pragmatic, epistemic/content, predictability/probabilistic, and provenance. Surface cues dominated the literature and were the most consistently operationalized. Discourse/pragmatic cues followed, particularly in discipline-bound academic genres where stance and metadiscourse differentiated AI from human writing. Predictability/probabilistic cues were central in detector-focused studies, while epistemic/content cues emerged primarily in tasks where grounding and authenticity were salient. Provenance cues were concentrated in watermarking research. Across studies, cue stability was consistently conditional rather than universal. Specifically, surface and discourse cues often remained discriminative within constrained genres, but shifted with register and discipline; probabilistic cues were powerful yet fragile under paraphrasing, post-editing, and evasion; and provenance signals required robustness to editing, mixing, and span localization. Overall, the literature indicates that AI–human distinction emerges from layered and context-dependent cue profiles rather than from any single reliable marker. High-stakes decisions, therefore, require condition-aware interpretation, triangulation across multiple cue families, and human oversight rather than automated classification in isolation.
Large language models (LLMs) are now routine writing tools across various domains, intensifying questions about when text should be treated as human-authored, artificial intelligence (AI)-generated, or collaboratively produced. This rapid review aimed to identify cue families reported in empirical studies as distinguishing AI from human-authored text and to assess how stable these cues are across genres/tasks, text lengths, and revision conditions. Following PRISMA guidelines, we searched four online databases for peer-reviewed English-language empirical articles (1 January 2022–1 January 2026). After deduplication and screening, 40 studies were included. Evidence converged on five cue families: surface, discourse/pragmatic, epistemic/content, predictability/probabilistic, and provenance cues. Surface cues dominated the literature and were the most consistently operationalized. Discourse/pragmatic cues followed, particularly in discipline-bound academic genres where stance and metadiscourse differentiated AI from human writing. Predictability/probabilistic cues were central in detector-focused studies, while epistemic/content cues emerged primarily in tasks where grounding and authenticity were salient. Provenance cues were concentrated in watermarking research. Across studies, cue stability was consistently conditional rather than universal. Specifically, surface and discourse cues often remained discriminative within constrained genres, but shifted with register and discipline; probabilistic cues were powerful yet fragile under paraphrasing, post-editing, and evasion; and provenance signals required robustness to editing, mixing, and span localization. Overall, the literature indicates that AI–human distinction emerges from layered and context-dependent cue profiles rather than from any single reliable marker. High-stakes decisions, therefore, require condition-aware interpretation, triangulation across multiple cue families, and human oversight rather than automated classification in isolation.
The rapid expansion of AI-generated writing has introduced significant challenges to academic integrity, particularly in relation to authorship verification within educational and research contexts. This study examines how AI-generated text can be distinguished from human-authored academic writing through a structured integration of data science methods, linguistic analysis, and insights drawn from existing student-centered research (Elkhatat et al., 2023; Opara, 2025; Weber-Wulff et al., 2023). Rather than proposing definitive detection outcomes, the study focuses on identifying recurring stylistic tendencies reported in prior work. The research adopts a mixed-methods review-oriented approach, combining quantitative stylometric analysis with qualitative textual interpretation. Stylometry, a well-established framework for analyzing writing style (Holmes, 1998; Stamatatos, 2009), is used to examine academic texts produced by multiple large language models—ChatGPT, Gemini, Claude, Grok, Perplexity, and DeepSeek—alongside essays written by undergraduate students, as documented in the reviewed literature. The analysis emphasizes observable linguistic features such as function word frequency, sentence structure regularity, and part-of-speech sequence patterns that reflect underlying stylistic behavior. Commonly reported stylometric markers include lexical diversity, average sentence length, syntactic dependency depth, punctuation usage, and recurrent n-gram patterns (Opara, 2025).
Generative artificial intelligence (GenAI), particularly large language model-based tools such as ChatGPT, has rapidly entered university English as a Foreign Language (EFL) writing instruction. These tools can support brainstorming, outlining, drafting, corrective feedback, revision, and academic language refinement. Yet their use also raises concerns about over-reliance, authorship, academic integrity, assessment validity, and the changing role of teachers in writing pedagogy. This systematic literature review synthesizes recent evidence on GenAI in university EFL writing instruction using the PRISMA 2020 framework. Searches were designed for Scopus, Web of Science Core Collection, ERIC, and Education Source/EBSCOhost, covering publications from 1 November 2022 to 22 June 2026. After duplicate removal, title/abstract screening, full-text eligibility assessment, and quality appraisal, 120 studies were included in the qualitative synthesis. Narrative thematic synthesis identified six recurring themes: GenAI as a writing-process scaffold, GenAI-generated feedback, revision uptake and learner engagement, teacher-AI feedback alignment, academic integrity and authorship, and methodological limitations in the existing evidence base. The review concludes that GenAI is most educationally defensible when integrated as a guided formative-feedback resource rather than as a substitute writer or replacement for teacher expertise. Practical implications are offered for assignment design, AI-use disclosure, feedback literacy, prompt literacy, and process-based assessment.