Machine learning algorithms can accurately detect anomalies in medical image scans.
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
11 sources for · 0 against
Multiple systematic reviews and clinical studies report that deep learning and machine learning algorithms achieve high diagnostic accuracy in detecting anomalies, lesions, and diseases across various medical imaging modalities.
Over the past decade, Deep Learning (DL) techniques have demonstrated remarkable advancements across various domains, driving their widespread adoption. Particularly in medical image analysis, DL received greater attention for tasks like image segmentation, object detection, and classification. This paper provides an overview of DL-based object recognition in medical images, exploring recent methods and emphasizing different imaging techniques and anatomical applications. Utilizing a meticulous quantitative and qualitative analysis following PRISMA guidelines, we examined publications based on citation rates to explore into the utilization of DL-based object detectors across imaging modalities and anatomical domains. Our findings reveal a consistent rise in the utilization of DL-based object detection models, indicating unexploited potential in medical image analysis. Predominantly within Medicine and Computer Science domains, research in this area is most active in the US, China, and Japan. Notably, DL-based object detection methods have gotten significant interest across diverse medical imaging modalities and anatomical domains. These methods have been applied to a range of techniques including CR scans, pathology images, and endoscopic imaging, showcasing their adaptability. Moreover, diverse anatomical applications, particularly in digital pathology and microscopy, have been explored. The analysis underscores the presence of varied datasets, often with significant discrepancies in size, with a notable percentage being labeled as private or internal, and with prospective studies in this field remaining scarce. Our review of existing trends in DL-based object detection in medical images offers insights for future research directions. The continuous evolution of DL algorithms highlighted in the literature underscores the dynamic nature of this field, emphasizing the need for ongoing research and fitted optimization for specific applications.
Background: Intracranial hemorrhage (ICH) is a life-threatening medical condition that needs early detection and treatment. In this systematic review and meta-analysis, we aimed to update our knowledge of the performance of deep learning (DL) models in detecting ICH on non-contrast computed tomography (NCCT). Methods: The study protocol was registered with PROSPERO (CRD420250654071). PubMed/MEDLINE and Google Scholar databases and the reference section of included studies were searched for eligible studies. The risk of bias in the included studies was assessed using the QUADAS-2 tool. Required data was collected to calculate pooled sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) with the corresponding 95% CI using the random effects model. Results: Seventy-three studies were included in our qualitative synthesis, and fifty-eight studies were selected for our meta-analysis. A pooled sensitivity of 0.92 (95% CI 0.90–0.94) and a pooled specificity of 0.94 (95% CI 0.92–0.95) were achieved. Pooled PPV was 0.84 (95% CI 0.78–0.89) and pooled NPV was 0.97 (95% CI 0.96–0.98). A bivariate model showed a pooled AUC of 0.96 (95% CI 0.95–0.97). Conclusions: This meta-analysis demonstrates that DL performs well in detecting ICH from NCCTs, highlighting a promising potential for the use of AI tools in various practice settings. More prospective studies are needed to confirm the potential clinical benefit of implementing DL-based tools and reveal the limitations of such tools for automated ICH detection and their impact on clinical workflow and outcomes of patients.
BACKGROUND AND OBJECTIVE
Magnetic resonance imaging (MRI) plays a critical role in prostate cancer diagnosis, but is limited by variability in interpretation and diagnostic accuracy. This systematic review evaluates the current state of deep learning (DL) models in enhancing the automatic detection, localization, and characterization of clinically significant prostate cancer (csPCa) on MRI.
METHODS
A systematic search was conducted across Medline/PubMed, Embase, Web of Science, and ScienceDirect for studies published between January 2020 and September 2023. Studies were included if these presented and validated fully automated DL models for csPCa detection on MRI, with pathology confirmation. Study quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) tool and the Checklist for Artificial Intelligence in Medical Imaging.
KEY FINDINGS AND LIMITATIONS
Twenty-five studies met the inclusion criteria, showing promising results in detecting and characterizing csPCa. However, significant heterogeneity in study designs, validation strategies, and datasets complicates direct comparisons. Only one-third of studies performed external validation, highlighting a critical gap in generalizability. The reliance on internal validation limits a broader application of these findings, and the lack of standardized methodologies hinders the integration of DL models into clinical practice.
CONCLUSIONS AND CLINICAL IMPLICATIONS
DL models demonstrate significant potential in improving prostate cancer diagnostics on MRI. However, challenges in validation, generalizability, and clinical implementation must be addressed. Future research should focus on standardizing methodologies, ensuring external validation and conducting prospective clinical trials to facilitate the adoption of artificial intelligence (AI) in routine clinical settings. These findings support the cautious integration of AI into clinical practice, with further studies needed to confirm their efficacy in diverse clinical environments.
PATIENT SUMMARY
In this study, we reviewed how artificial intelligence (AI) models can help doctors better detect and understand aggressive prostate cancer using magnetic resonance imaging scans. We found that while these AI tools show promise, these tools need more testing and validation in different hospitals before these can be used widely in patient care.
The semantic segmentation (SS) of low-contrast images (LCIs) remains a significant challenge in computer vision, particularly for sensor-driven applications like medical imaging, autonomous navigation, and industrial defect detection, where accurate object delineation is critical. This systematic review develops a comprehensive evaluation of state-of-the-art deep learning (DL) techniques to improve segmentation accuracy in LCI scenarios by addressing key challenges such as diffuse boundaries and regions with similar pixel intensities. It tackles primary challenges, such as diffuse boundaries and regions with similar pixel intensities, which limit conventional methods. Key advancements include attention mechanisms, multi-scale feature extraction, and hybrid architectures combining Convolutional Neural Networks (CNNs) with Vision Transformers (ViTs), which expand the Effective Receptive Field (ERF), improve feature representation, and optimize information flow. We compare the performance of 25 models, evaluating accuracy (e.g., mean Intersection over Union (mIoU), Dice Similarity Coefficient (DSC)), computational efficiency, and robustness across benchmark datasets relevant to automation and robotics. This review identifies limitations, including the scarcity of diverse, annotated LCI datasets and the high computational demands of transformer-based models. Future opportunities emphasize lightweight architectures, advanced data augmentation, integration with multimodal sensor data (e.g., LiDAR, thermal imaging), and ethically transparent AI to build trust in automation systems. This work contributes a practical guide for enhancing LCI segmentation, improving mean accuracy metrics like mIoU by up to 15% in sensor-based applications, as evidenced by benchmark comparisons. It serves as a concise, comprehensive guide for researchers and practitioners advancing DL-based LCI segmentation in real-world sensor applications.
Abstract Background Thyroid cancer is one of the most common endocrine malignancies. Its incidence has steadily increased in recent years. Distinguishing between benign and malignant thyroid nodules (TNs) is challenging due to their overlapping imaging features. The rapid advancement of artificial intelligence (AI) in medical image analysis, particularly deep learning (DL) algorithms, has provided novel solutions for automated TN detection. However, existing studies exhibit substantial heterogeneity in diagnostic performance. Furthermore, no systematic evidence-based research comprehensively assesses the diagnostic performance of DL models in this field. Objective This study aimed to execute a systematic review and meta-analysis to appraise the performance of DL algorithms in diagnosing TN malignancy, identify key factors influencing their diagnostic efficacy, and compare their accuracy with that of clinicians in image-based diagnosis. Methods We systematically searched multiple databases, including PubMed, Cochrane, Embase, Web of Science, and IEEE, and identified 41 eligible studies for systematic review and meta-analysis. Based on the task type, studies were categorized into segmentation (n=14) and detection (n=27) tasks. The pooled sensitivity, specificity, and the area under the receiver operating characteristic curve (AUC) were calculated for each group. Subgroup analyses were performed to examine the impact of transfer learning and compare model performance against clinicians. Results For segmentation tasks, the pooled sensitivity, specificity, and AUC were 82% (95% CI 79%‐84%), 95% (95% CI 92%‐96%), and 0.91 (95% CI 0.89‐0.94), respectively. For detection tasks, the pooled sensitivity, specificity, and AUC were 91% (95% CI 89%‐93%), 89% (95% CI 86%‐91%), and 0.96 (95% CI 0.93‐0.97), respectively. Some studies demonstrated that DL models could achieve diagnostic performance comparable with, or even exceeding, that of clinicians in certain scenarios. The application of transfer learning contributed to improved model performance. Conclusions DL algorithms exhibit promising diagnostic accuracy in TN imaging, highlighting their potential as auxiliary diagnostic tools. However, current studies are limited by suboptimal methodological design, inconsistent image quality across datasets, and insufficient external validation, which may introduce bias. Future research should enhance methodological standardization, improve model interpretability, and promote transparent reporting to facilitate the sustainable clinical translation of DL-based solutions.
<h4>Topic</h4>To evaluate the performance of machine learning (ML) in the diagnosis of retinopathy of prematurity (ROP) and to assess whether it can be an effective automated diagnostic tool for clinical applications.<h4>Clinical relevance</h4>Early detection of ROP is crucial for preventing tractional retinal detachment and blindness in preterm infants, which has significant clinical relevance.<h4>Methods</h4>Web of Science, PubMed, Embase, IEEE Xplore, and Cochrane Library were searched for published studies on image-based ML for diagnosis of ROP or classification of clinical subtypes from inception to October 1, 2022. The quality assessment tool for artificial intelligence-centered diagnostic test accuracy studies was used to determine the risk of bias (RoB) of the included original studies. A bivariate mixed effects model was used for quantitative analysis of the data, and the Deek's test was used for calculating publication bias. Quality of evidence was assessed using Grading of Recommendations Assessment, Development and Evaluation.<h4>Results</h4>Twenty-two studies were included in the systematic review; 4 studies had high or unclear RoB. In the area of indicator test items, only 2 studies had high or unclear RoB because they did not establish predefined thresholds. In the area of reference standards, 3 studies had high or unclear RoB. Regarding applicability, only 1 study was considered to have high or unclear applicability in terms of patient selection. The sensitivity and specificity of image-based ML for the diagnosis of ROP were 93% (95% confidence interval [CI]: 0.90-0.94) and 95% (95% CI: 0.94-0.97), respectively. The area under the receiver operating characteristic curve (AUC) was 0.98 (95% CI: 0.97-0.99). For the classification of clinical subtypes of ROP, the sensitivity and specificity were 93% (95% CI: 0.89-0.96) and 93% (95% CI: 0.89-0.95), respectively, and the AUC was 0.97 (95% CI: 0.96-0.98). The classification results were highly similar to those of clinical experts (Spearman's R = 0.879).<h4>Conclusions</h4>Machine learning algorithms are no less accurate than human experts and hold considerable potential as automated diagnostic tools for ROP. However, given the quality and high heterogeneity of the available evidence, these algorithms should be considered as supplementary tools to assist clinicians in diagnosing ROP.<h4>Financial disclosure(s)</h4>Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Deep learning (DL) has revolutionized medical image analysis (MIA), enabling early anomaly detection, precise lesion segmentation, and automated disease classification. However, its clinical integration faces two major challenges: reliance on limited, narrowly annotated datasets that inadequately capture real-world patient diversity, and the inherent “black-box” nature of DL decision-making, which complicates physician scrutiny and accountability. Eye tracking (ET) technology offers a transformative solution by capturing radiologists’ gaze patterns to generate supervisory signals. These signals enhance DL models through two key mechanisms: providing weak supervision to improve feature recognition and diagnostic accuracy, particularly when labeled data are scarce, and enabling direct comparison between machine and human attention to bridge interpretability gaps and build clinician trust. This approach also extends effectively to multimodal learning models (MLMs) and vision–language models (VLMs), supporting the alignment of machine reasoning with clinical expertise by grounding visual observations in diagnostic context, refining attention mechanisms, and validating complex decision pathways. Conducted in accordance with the PRISMA statement and registered in PROSPERO (ID: CRD42024569630), this review synthesizes state-of-the-art strategies for ET-DL integration. We further propose a unified framework in which ET innovatively serves as a data efficiency optimizer, a model interpretability validator, and a multimodal alignment supervisor. This framework paves the way for clinician-centered AI systems that prioritize verifiable reasoning, seamless workflow integration, and intelligible performance, thereby addressing key implementation barriers and outlining a path for future clinical deployment.
Medical imaging abnormality detection is challenging, but deep learning approaches have shown promise. This paper reviews the current state of the art in deep learning approaches for detecting abnormalities in chest medical imaging. To discover the trends, opportunities, and challenges associated with this field, 18 studies were selected from Google Scholar based on their titles, abstracts, and contents for extensive review to answer two research questions. The study found that the National Institutes of Health (NIH) Chest X-ray 14 dataset is the most used dataset for this task. Most research uses a single-modal approach, considering only image data as input, with X-ray being the more popular instrument. There are 8 out of 18 studies leverage the transfer learning approach, with ResN et50 being the most popular network. MobileNetV2 has demonstrated competitive results compared to more robust networks. Preprocessing techniques such as image enhancement and data augmentation are leveraged by 61.1 % of the reviewed studies and are shown to improve model performance.
Machine learning algorithms have the potential to revolutionize the way healthcare providers care for pregnant women. By using various input variables such as maternal characteristics, medical history, ultrasound measurements, and biomarkers, machine learning algorithms can develop personalized risk assessments to guide clinical decision-making and interventions. Additionally, these algorithms can monitor fetal well-being by analyzing electronic fetal monitoring data and detect signs of fetal distress. Image analysis algorithms can also identify fetal anomalies or complications more accurately and efficiently than manual interpretation of images. Finally, machine learning algorithms can develop personalized treatment recommendations based on clinical data, such as identifying optimal medication dosages and recommending the most appropriate delivery mode for women with prior cesarean deliveries. Overall, leveraging machine learning can improve the care of pregnant women and help ensure healthy outcomes for both mother and baby.
2025 meta-analysis in PLOS One found that the use of AI algorithms for detecting tooth decay was clinically justified. AI algorithms have been created
Artificial intelligence in healthcare refers to the application of artificial intelligence (AI) to medical and healthcare data in areas including disease diagnosis, treatment planning,, patient monitoring, drug development, and clinical decision support systems.
The use of AI in healthcare has also raised ethical, technical, and regulatory concerns, including related to issues such as data privacy
Intel's venture capital arm Intel Capital invested in 2016 in the startup Lumiata, which uses AI to identify at-risk patients and develop care options.
Siemens Healthineers applies AI in imaging and diagnostics, including algorithms to reconstruct CT images and guide ultrasound procedures. It also uses AI to support treatment planning such as radiation therapy for cancer, improve point-of-care diagnostics, and automate lab workflows.
Microsoft's Hanover project, in partnership with Oregon Health & Science University's Knight Cancer Institute, analyzes medical research to predict the most effective cancer drug treatment options for patients. Other projects include medical image analysis of tumor progression and the development of programmable cells.
Mayo Clinic, in partnership with Microsoft, is designing a model centered around healthcare, aimed at handling clinical reasoning tasks, supporting earlier diagnoses and personalized treatment decisions.
Philips Healthcare develops AI-powered diagnostic tools that analyze medical images to detect subtle anomalies. Its AI technologies also support precision oncology by assisting pathologists in cancer diagnosis, care management, and patient monitoring.
Tencent has been working on several medical systems and services. These include AI Medical Innovation System (AIMIS), an AI-powered diagnostic medical imaging service; WeChat Intelligent Healthcare; and Tencent Doctorwork
A comparison of two approaches to three-dimensional imaging of craniofacial anomalies.
Volume-based and surface-based algorithms for three-dimensional rendering of computed tomography (CT) scans of the human skull were compared in patients with craniofacial anomalies. Both methods were applied to a selected sample of 12 clinical CT studies. The number of sections ranged from 24 to 72 and the section thickness from 1.5 to 6.0 mm. Volume renderings were more prone to interpolation artifacts but captured the anatomy in greater detail. The sites of closed cranial sutures, visualized using the volume technique, were not demonstrated using the specific surface rendering technique used in this study. In both techniques the areas of thin bone appeared as gaps.
Published in Journal of digital imaging (1990)
Everything we examined (11)
This check searched the claim as stated. It did not run a separate search for evidence against it.