Natural language understanding can be scientifically defined and measured through standardized tests
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
8 sources for · 0 against
Natural language understanding in computational and artificial intelligence systems is widely defined and measured through various standardized benchmark datasets and tests evaluating specific linguistic and reasoning tasks.
In light of recent breakthroughs in large language models (LLMs) that have revolutionized natural language processing (NLP), there is an urgent need for new benchmarks to keep pace with the fast development of LLMs. In this paper, we propose CFLUE, the Chinese Financial Language Understanding Evaluation benchmark, designed to assess the capability of LLMs across various dimensions. Specifically, CFLUE provides datasets tailored for both knowledge assessment and application assessment. In knowledge assessment, it consists of 38K+ multiple-choice questions with associated solution explanations. These questions serve dual purposes: answer prediction and question reasoning. In application assessment, CFLUE features 16K+ test instances across distinct groups of NLP tasks such as text classification, machine translation, relation extraction, reading comprehension, and text generation. Upon CFLUE, we conduct a thorough evaluation of representative LLMs. The results reveal that only GPT-4 and GPT-4-turbo achieve an accuracy exceeding 60\% in answer prediction for knowledge assessment, suggesting that there is still substantial room for improvement in current LLMs. In application assessment, although GPT-4 and GPT-4-turbo are the top two performers, their considerable advantage over lightweight LLMs is noticeably diminished. The datasets and scripts associated with CFLUE are openly accessible at https://github.com/aliyun/cflue.
[2405.10542] Benchmarking Large Language Models on CFLUE - A Chinese Financial Language Understanding Evaluation Dataset Benchmarking Large Language Models on CFLUE - A Chinese Financial Language Understanding Evaluation Dataset Jie Zhu 1 , Junhui Li 2 , Yalong Wen 1 , Lifan Guo 1 1 Alibaba Group, Hangzhou, China 2 School of Computer Science and Technology, Soochow University, Suzhou, China zhujie951121@gmail.com, lijunhui@suda.edu.cn {wenyalong.wyl, lifan.lg}@alibaba-inc.com Corresponding Author Abstract In light of recent breakthroughs in large language models (LLMs) that have revolutionized natural language processing (NLP), there is an urgent need
In this paper, we propose CFLUE, the Chinese Financial Language Understanding Evaluation benchmark, designed to assess the capability of LLMs across various dimensions. Specifically, CFLUE provides datasets tailored for both knowledge assessment and application assessment. In knowledge assessment, it consists of 38K+ multiple-choice questions with associated solution explanations. These questions serve dual purposes: answer prediction and question reasoning. In application assessment, CFLUE features 16K+ test instances across distinct groups of NLP tasks such as text classification, machine translation, relation extraction, reading comprehension, and text generation.
Benchmarking Large Language Models on CFLUE - A Chinese Financial Language Understanding Evaluation Dataset Jie Zhu 1 , Junhui Li 2 † † thanks: Corresponding Author , Yalong Wen 1 , Lifan Guo 1 1 Alibaba Group, Hangzhou, China 2 School of Computer Science and Technology, Soochow University, Suzhou, China zhujie951121@gmail.com, lijunhui@suda.edu.cn {wenyalong.wyl, lifan.lg}@alibaba-inc.com 1 Introduction Recently, the remarkable capabilities exhibited by large language models (LLMs) have brought significant advancements and revolutionized natural language processing (NLP).
Additionally, existing shared tasks, such as those in CCKS Tianchi ( 2019 , 2020 , 2021 , 2022 ) , predominantly concentrate on event extraction tasks, limiting the objective and quantitative measurement of LLM performance. Inspired by the FLUE benchmark Shah et al. ( 2022 ) , which encompasses a comprehensive set of datasets across five financial domain tasks in English, this paper introduces a novel dataset named CFLUE (Chinese Financial Language Understanding Evaluation). CFLUE addresses the aforementioned challenges by providing a benchmark for evaluating LLM performance through various NLP tasks, categorized into knowledge assessment and application assessment.
Additionally, beyond evaluation suites, there are financial datasets like SmoothNLP 2 2 2 https://github.com/smoothnlp/FinancialDatasets , IREE Ren et al. ( 2022 ) , suitable for training or fine-tuning models in the finance domain. Other Benchmark Datasets. The development of LMs Devlin et al. ( 2019 ); Radford et al. ( 2019 ) has witnessed heterogeneous benchmarks to probe their diverse abilities. In English, conventional benchmarks traditionally target single (type) tasks such as natural language understanding Wang et al. ( 2019b , a ) , reading comprehension Rajpurkar et al. ( 2018 ); Dua et al. ( 2019 ) , and reasoning Zellers et al. ( 2019 ); Sakaguchi et al. ( 2021 ) , etc.
3 CFLUE: Chinese Financial Language Understanding Evaluation Figure 1: Overview diagram of CFLUE benchmark. Type Task Size Avg. Length Description Train Valid Test Knowl. Answer Prediction & Reasoning 30,908 3,864 3,864 130 15 types of qualification exams Appl. Text Classification - - 3,312 840 6 subtasks Machine Translation - - 3,000 51 2 translation directions Relation Extraction - - 3,500 274 4 subtasks Reading Comprehension - - 2,710 201 in question answering format Text Generation - - 4,000 947 5 subtasks Total - - 16,522 - - Table 2: Statistics of CFLUE. The average length indicates the average number of words in the inputs.
Llama: Open and efficient foundation language models. Computing Research Repository , arXiv:2302.13971. Wang et al. (2019a) Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a. Superglue: A stickier benchmark for general-purpose language understanding systems. In Proceedings of NeurIPS , pages 3266–3280. Wang et al. (2019b) Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019b. Glue: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of ICLR . Xu et al.
Large language models (LLMs) can understand natural language and generate corresponding text, images, and even videos based on prompts, which holds great potential in medical scenarios. Orthopedics is a significant branch of medicine, and orthopedic diseases contribute to a significant socioeconomic burden, which could be alleviated by the application of LLMs. Several pioneers in orthopedics have conducted research on LLMs across various subspecialties to explore their performance in addressing different issues. However, there are currently few reviews and summaries of these studies, and a systematic summary of existing research is absent. The objective of this review was to comprehensively summarize research findings on the application of LLMs in the field of orthopedics and explore the potential opportunities and challenges. PubMed, Embase, and Cochrane Library databases were searched from January 1, 2014, to February 22, 2024, with the language limited to English. The terms, which included variants of "large language model," "generative artificial intelligence," "ChatGPT," and "orthopaedics," were divided into 2 categories: large language model and orthopedics. After completing the search, the study selection process was conducted according to the inclusion and exclusion criteria. The quality of the included studies was assessed using the revised Cochrane risk-of-bias tool for randomized trials and CONSORT-AI (Consolidated Standards of Reporting Trials-Artificial Intelligence) guidance. Data extraction and synthesis were conducted after the quality assessment. A total of 68 studies were selected. The application of LLMs in orthopedics involved the fields of clinical practice, education, research, and management. Of these 68 studies, 47 (69%) focused on clinical practice, 12 (18%) addressed orthopedic education, 8 (12%) were related to scientific research, and 1 (1%) pertained to the field of management. Of the 68 studies, only 8 (12%) recruited patients, and only
Background Large language models (LLMs) can understand natural language and generate corresponding text, images, and even videos based on prompts, which holds great potential in medical scenarios. Orthopedics is a significant branch of medicine, and orthopedic diseases contribute to a significant socioeconomic burden, which could be alleviated by the application of LLMs. Several pioneers in orthopedics have conducted research on LLMs across various subspecialties to explore their performance in addressing different issues. However, there are currently few reviews and summaries of these studies, and a systematic summary of existing research is absent.
In recent years, this area has emerged as one of the most prominent areas of research in artificial intelligence (AI) innovation [ 1 , 2 ]. What makes LLMs different from smaller-scale PLMs is their remarkable emergent abilities to solve complex tasks. Studies have found that LLMs, such as generative pretrained transformer (GPT)-3 with approximately 175 billion parameters, exhibit a significant leap in natural language processing (NLP) capabilities compared to PLMs with fewer parameters, such as GPT-2 with approximately 1.5 billion parameters [ 2 , 3 ].
Generative AI applications developed based on LLMs not only possess the ability to understand natural language but can also generate corresponding text, images, and even videos based on input sources. This human-machine interaction mode holds great potential in medical scenarios. LLMs have undergone significant advancements in recent years; currently, the most prevalent web-based LLM service is ChatGPT (OpenAI). Launched in November 2022, ChatGPT is a chatbot application developed based on GPT-3.5 or GPT-4 after fine-tuning, and it can quickly respond to questions posed by users.
The inclusion and exclusion criteria are listed in Textbox 2 . The results were cross-checked, and discrepancies were resolved through discussion, with the final determination made by a third investigator (YT). Inclusion and exclusion criteria.
The feasibility of this approach has been validated in interdisciplinary research on the management of back pain [ 81 ]. Application of LLMs in Management Trained NLP models can convert natural language into structured data and have demonstrated superior performance in tasks involving the current procedural terminology for identifying spinal surgery records [ 93 ]. However, ChatGPT, with its larger parameters, performs weaker than NLP models in the task of identifying spinal surgery current procedural terminology codes [ 85 ].
The potential reasons for the lack of readability may include not only the limited training data but also the quality of the trained data. By incorporating more popular science content and common clinical responses, it may be possible to address the issue of readability through fine-tuning the model. Reliability Different ways of asking the same question may yield completely different answers [ 21 ]. This instability, particularly in response to specific prompts, not only affects users’ experience and trust but also greatly interferes with researchers’ homogenized evaluations. It is imperative to establish standardized questioning processes and prompt criteria.
Furthermore, the absence of standardized questioning paradigms has led to instability in LLM responses, posing challenges for reproducibility and limiting the reliability and clinical significance of the study findings. Limitations of This Review This systematic review has several limitations. First, only English-language articles were included, which may have led to the exclusion of relevant studies published in other languages. Second, due to significant heterogeneity in study designs, model tasks, and evaluation parameters among the included studies, we did not perform a comprehensive synthesis of most of the data, nor did we conduct a meta-analysis.
include learning, reasoning, knowledge representation, planning, natural language processing, and perception, as well as support for robotics. To reach these
Artificial intelligence (AI) is the capability of computational systems to perform tasks typically associated with human intelligence, such as learning, reasoning, problem-solving, perception, and decision-making. It is a field of research in engineering, mathematics, and computer science that develops and studies methods and software that enable machines to perceive their environment and use lear
Artificial intelligence (AI) is the capability of computational systems to perform tasks typically associated with human intelligence, such as learning, reasoning, problem-solving, perception, and decision-making. It is a field of research in engineering, mathematics, and computer science that develops and studies methods and software that enable machines to perceive their environment and use learning and intelligence to take actions that maximize their chances of achieving defined goals.
High-profile applications of AI include advanced web search engines, chatbots, virtual assistants, autonomous vehicles, play and analysis in strategy games (e.g., chess and Go), and content generation (e.g. images, audio, and videos).
The traditional goals of AI research include learning, reasoning, knowledge representation, planning, natural language processing, and perception, as well as support for robotics. To reach these goals, AI researchers use techniques including state space search and mathematical optimization, formal logic, artificial neural networks, and methods based on statistics, operations research, and economics. AI also draws upon psychology, linguistics, philosophy, neuroscience, and other fields. Some companies, such as OpenAI, Google DeepMind, and Meta, aim to create artificial general intelligence (AGI)—AI that can complete nearly any cognitive task at least as well as a human.
Artificial intelligence was founded as an academic discipline in 1956. The field went through multiple cycles of optimism throughout its history, followed by periods of disappointment and loss of funding, known as AI winters. Funding and interest increased substantially after 2012, when graphics processing units (GPUs) started being used to accelerate neural networks, and deep learning outperformed previous AI techniques. This growth accelerated further after 2017 with the transformer architecture. In the 2020s, an AI boom coincided with advances in generative AI, which became widespread and allowed for the creation and modification of media. In addition to AI safety and unintended consequences and harms from the use of AI, ethical concerns, AI's long-term effects, environmental…
<h4>Introduction</h4>Large language models (LLMs), including OpenAI's GPT family accessed via interfaces such as ChatGPT and Microsoft Copilot, as well as non-GPT systems such as Google Gemini, are increasingly applied in healthcare and dental education. However, the accuracy of these systems in specialized tasks such as answering dental examination questions remains unclear.<h4>Methods</h4>This systematic review and meta-analysis evaluated LLM performance in answering dental questions. Databases searched were PubMed, Embase, Scopus, and Web of Science. Data on question type and number, LLM versions, and accuracy rates were extracted. Pooled accuracy was estimated using a random-effects model; heterogeneity and publication bias were assessed.<h4>Results</h4>A total of 39 studies were included, with ChatGPT-4 being the most frequently evaluated model. The pooled accuracy for LLMs was 63.7% (95% CI: 60.3%-67.1%), with high heterogeneity (I² = 91.5%). Subgroup analysis revealed ChatGPT-4 and Copilot (a GPT-based interface) achieved the highest pooled accuracies (∼73% and ∼75%, respectively). Direct comparisons confirmed ChatGPT-4 significantly outperformed earlier versions and some competitor models. Sensitivity analyses supported the robustness of findings.<h4>Conclusion</h4>LLMs demonstrate moderate accuracy in answering dental examination questions and are currently insufficient for autonomous clinical decision-making. When their limitations are explicitly recognized, however, these systems may serve as valuable adjuncts in dental education and examination preparation. Methodological strategies such as structured prompting and retrieval-augmented approaches warrant further investigation but were not the primary focus of the present analysis.
of standardized educational and mental tests . As we shall see later, his scores on the various intelligence scales ranged from 12 to 13½ years, and on
information system can be significantly and economically extended through direct communication between users and the system in natural language . Unfortunately
Everything we examined (8) — 7 independent sources
This check searched the claim as stated. It did not run a separate search for evidence against it.